microFlare v1

microFlare v1 is an asymmetrically quantized version of Qwen 3.6 35B-A3B. The model has been optimized layer by layer, with high precision for important parts and lower bitrates for less consequential weights. It is designed so it can be ran on low spec hardware, as low as a 4GB GPU + 16GB of system RAM utilizing CPU offloading of experts. It should hopefully fit well onto a 16GB GPU without need for CPU offload, but this has not yet been tested. Additionally if you don't even have a 4GB GPU, it can be ran completely on CPU with 16GB of system RAM, but it will be slow probably.

Note that the MTP (layer 64) predictor has been stripped out of this version, as it proved to be detrimental on older hardware. I am considering releasing another version with the MTP included, or maybe just a stand alone file for it (let me know which would be more useful to you).

Perplexity drift

Ran with a 2k context, revision A has a relative perplexity score* of 6.2761 +/- 0.03961
*note scoring on my machine is different than as reported elsewhere, so this is better than it may sound at first

Model Size Perplexity Score
bartowski/Q4_K 22GB 5.9654 +/- 0.03726
unsloth/UD-IQ3_S 13.7GB 6.2496 +/- 0.03948
microFlare v1.a 14GB 6.2761 +/- 0.03961
mudler/APEX-I-Mini 14.3GB 6.3554 +/- 0.04057
bartowski/Q2_K_L 14GB 6.4180 +/- 0.04066

Test Suites

GSM8K (basic math)

Ran with 5 shot and 2048 context

Match type Score
flexible 0.5906 ยฑ0.0135
strict 0.6058 ยฑ0.0135

MMLU Pro (comprehensive)*

*note these results are from limited partial run, only the first 20 questions of the 14 categories were tested

Ran with 5 shot and 8196 context

Tasks Version Filter n-shot Metric Value Stderr
Total Average 2.0 custom-extract 5 exact_match 0.7286 ยฑ0.0262
- biology 3.1 custom-extract 5 exact_match 0.8000 ยฑ0.0918
- business 3.1 custom-extract 5 exact_match 0.7500 ยฑ0.0993
- chemistry 3.1 custom-extract 5 exact_match 0.8500 ยฑ0.0819
- computer_science 3.1 custom-extract 5 exact_match 0.8500 ยฑ0.0819
- economics 3.1 custom-extract 5 exact_match 0.6500 ยฑ0.1094
- engineering 3.1 custom-extract 5 exact_match 0.5500 ยฑ0.1141
- health 3.1 custom-extract 5 exact_match 0.7500 ยฑ0.0993
- history 3.1 custom-extract 5 exact_match 0.7000 ยฑ0.1051
- law 3.1 custom-extract 5 exact_match 0.4500 ยฑ0.1141
- math 3.1 custom-extract 5 exact_match 0.9000 ยฑ0.0688
- other 3.1 custom-extract 5 exact_match 0.6000 ยฑ0.1124
- philosophy 3.1 custom-extract 5 exact_match 0.8000 ยฑ0.0918
- physics 3.1 custom-extract 5 exact_match 0.7000 ยฑ0.1051
- psychology 3.1 custom-extract 5 exact_match 0.8500 ยฑ0.0819

Manual prompt testing

So far all testing has been successful. Of the 50 some odd prompts it's been tested with everyone came back with a high quality answer/result. It built a pretty decent test website on a one shot, so coding abilities are intact. It has not been tested with any tools or agent harnesses yet, nor image based inputs. I've only tested text so far.

For more details and tips on running optimally, see my full blog post.

Downloads last month
68
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for MicroFlare/microFlare-v1

Quantized
(745)
this model