First ; thank you!
Based on early checking/Qwen's own statements (uses same Qwen 3.5 arch as 3.5,.6) -> it is doable.
But need to verify everything.
First ; thank you!
Based on early checking/Qwen's own statements (uses same Qwen 3.5 arch as 3.5,.6) -> it is doable.
But need to verify everything.
RE: deepseek V4 ; sorry don't have the VRAM to do it.
Yes. It is on the list.
We are still revising the 9-14B pipeline.
We do have an interm Qwen 3.5 9B from the experimental pipeline here:
It matches/exceeds 27B Qwen 3.5 ; and meets in some cases 27B Qwen 3.6 performance.
There is still a lot of optimizations to do at this time.
Try the Q6, and/or IQ4_XS.
Even IQ3_M will be very strong.
Try this one:
https://huggingface.co/DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF
This is the "smaller version" of the Fable Fusion 711.
It operates at almost 27B power at only 9B parameters.
This will far exceed Qwen 2.5 and Qwen 3 performance.
Excellent. The MOE variants require a lot more VRAM/time for training.
Please take a moment to visit the repo where some of your concerns are addressed on the repo card itself.
2nd; publishing all the metrics at each step would be both exhausting and worse confusing.
I don't follow what you mean by "cheap" ; as a heretic [step] you usually lose 2-4 points on some metrics.
So the comparison of "heretic" vs "non-heretic" is even STRONGER ; the fairer one would be "heretic base" to "heretic tuned" which would likely show even greater change / improvement.
Source/MLX here:
https://huggingface.co/nightmedia/Qwen3.5-9B-DS9-USS-Defiant
This is on my partner's repo.
Here is a new one just uploaded; also off the scale strong, but at 9B:
Benches are up ; beats Qwen 3.5 27B in all 7 benches AND Qwen3.6 35B-A3B.
Matches some Qwen 3.6 27B benches too.
Clocks in at over 640 ARC-C for both 8bit and 4bit.
1/3 the size almost all the firepower.
Currently waiting on finalization of "4 bit" compression for these model types to address tuning/Vram issues.
These are in progress at BNB / Unsloth.
That is really the only hold up.
Otherwise VRAM to train these sparse moes is in 70 to 100 GB range. And really slow too.
Sorry no, not at this time.
This model does not contain MTP layers ; you need to run at non-MTP.
As of this writing:
There are pipeline (issues as well as optimizations) issues still currently, and it is not widely supported in some AI Apps.
Specifically:
Ggufs:
Training is compounded by number of experts in the model, which adds a serious level of time to the training.
Even 1000 samples [small!] takes 6-12 hrs.
Consider 31B dense , same samples, 30-60 minutes.
I will add to the list; may wait for specific Heretic and/or tuned version.
I already have a 43B-A3B version running in the lab ; however tuning these sparse moe models take a lot more work/time and ahh... detail. AND a lot more VRAM!!! [can't compress these atm, so BF16 required => 100 GB+ ]