AI & ML interests

PPALLI

SeaWolf-AIΒ 
posted an update 7 days ago
view post
Post
6348
πŸ–ΌοΈ POCKET-Image β€” the POCKET series goes visual: character-perfect text in any language, on-device

A new model in VIDRAFT's POCKET family. POCKET put 35B-class models on phones and no-GPU PCs. POCKET-Image carries the same "big capability, small hardware" idea into image generation β€” and fixes the one thing nearly every image model gets wrong: text.

Type "μ•ˆλ…•ν•˜μ„Έμš”" into a typical model and you get "μ•ˆγ…κΈ°." Hangul alone composes 11,172 syllable blocks; Arabic connects its letters; Thai stacks marks. Diffusion models draw scripts as shapes, so they smear. POCKET-Image renders every glyph exactly β€” ν•œκ΅­μ–΄ Β· δΈ­ζ–‡ Β· ζ—₯本θͺž Β· Ψ§Ω„ΨΉΨ±Ψ¨ΩŠΨ© (RTL) Β· ΰΉ„ΰΈ—ΰΈ’ Β· Latin and more β€” onto any scene you describe.

What it is:
β€’ 100% accurate text, any language β€” where global models produce gibberish
β€’ Any background from a prompt β€” text is optional (empty β†’ a pure image)
β€’ No GPU, no NPU β€” runs on plain CPU + RAM via the POCKET-Core engine
β€’ Measured footprint: 8.6 GB (RTX 3050/4060) Β· 4.5 GB (offloaded, 6 GB cards) Β· 13.4 GB (MacBook, 16 GB+)
β€’ Windows Β· macOS Β· Linux Β· fully local, no cloud

Built on the open, commercial-friendly Z-Image (Apache-2.0) foundation.

Honest note: the text is the guaranteed-correct part β€” the surrounding scene is ordinary generation, so a busy foreground can crowd the letters. We say so; clean backgrounds stay razor-sharp.

🎨 Studio β€” generate right here, any language:
FINAL-Bench/POCKET-Image-Studio

🧩 Model card:
FINAL-Bench/POCKET-Image-Zimage

πŸ“š The POCKET collection:
https://huggingface.co/collections/FINAL-Bench/pocket-models
SeaWolf-AIΒ 
posted an update 8 days ago
view post
Post
3115
POCKET now speaks Gemma 4 β€” a 26B model that loads in every app, and runs on your PC with no GPU

We're adding a Gemma-4 sibling to POCKET: POCKET-26B, built from Google's Gemma-4-26B-A4B (Apache-2.0). Our flagship POCKET-35B is a Qwen-family MoE and needs a recent llama.cpp; POCKET-26B trades a little size for the thing people kept asking for β€” it just loads, everywhere, today: Ollama, LM Studio, PocketPal, MLX, any stock llama.cpp. No fork, no bleeding-edge runtime, no CUDA, no cloud.

It's a sparse Mixture-of-Experts (25.2B total, ~4B active per token), so the work per token stays small β€” a real 26B that generates on a CPU with no graphics card.

Two things make it stand out:

1) Universal compatibility. Gemma 4 is a standard, widely-supported architecture, so POCKET-26B runs on the tools you already have β€” no waiting for your app to add a new model type.

2) Quality that survives compression. Measured GPQA-Diamond (198 q, greedy):
β€’ Full base: 67.7%
β€’ POCKET-26B Q4_K_M (17 GB): 67.7% β€” lossless
β€’ POCKET-26B Q2_K (11 GB): 67.2% β€” near-lossless, at 11 GB

Live, on a CPU-only box (our demo Space β€” POCKET-26B vs Bonsai-27B, same machine, same stock llama.cpp): POCKET-26B β‰ˆ 19 tok/s vs Bonsai β‰ˆ 6 tok/s β†’ about 3Γ— faster generation, no GPU. (Honest notes: shared CPU box, sequential race; a dedicated machine is faster.)

Where it fits in the family:
β€’ POCKET-35B (Qwen MoE) β€” bigger, top-tier, needs a recent llama.cpp.
β€’ POCKET-26B (Gemma 4) β€” loads in any app, quality-robust when compressed. The demo runs the Q4_K_M build; Q2_K (11 GB) is the smallest footprint. For a true ≀8 GB phone, the 5 GB POCKET-KR (Qwen) is still the pick.

Try it and grab it:
πŸ–₯️ Live demo (Gemma4-based, answering on a CPU, no GPU): FINAL-Bench/POCKET-26B-CPU
πŸ“¦ POCKET-26B-GGUF (Q4_K_M 17 GB Β· Q2_K 11 GB): FINAL-Bench/POCKET-26B-GGUF
πŸ“š POCKET collection: https://huggingface.co/collections/FINAL-Bench/pocket-models
  • 3 replies
Β·
SeaWolf-AIΒ 
posted an update 10 days ago
view post
Post
4023
πŸ“± POCKET β€” a 35-billion-parameter model that runs on your iPhone, and on your PC with no GPU

We're releasing POCKET, VIDRAFT's flagship Darwin-36B-Opus compressed for on-device use. No fork, no CUDA, no cloud β€” it runs on stock llama.cpp. It's a sparse Mixture-of-Experts model (256 experts, only 8 active per token), so the file can be large while the work per token stays small. That's what lets a 35B model run on a phone, and generate fast on a CPU with no graphics card.

Measured (POCKET-35B IQ1_M vs Bonsai-27B Q1_0):
β€’ CPU generate (Xeon, 16 threads): 27.0 vs 10.1 tok/s β†’ 2.69Γ— faster
β€’ GPU generate (H100): 197 vs 89 tok/s β†’ 2.22Γ— faster
β€’ GPU prompt processing (H100): 753 vs 1816 β†’ 0.41Γ— (Bonsai wins this one β€” MoE prefill wakes every expert, so sparsity stops helping there. We say so.)
β€’ Quality (HellaSwag, 400 q): 61.0% vs 60.0% β†’ a tie (confidence intervals overlap)

On a real consumer laptop β€” MacBook M3 Pro (18 GB) β€” POCKET wins every axis, prompt processing included:
β€’ Metal generate: 25.4 vs 12.8 β†’ 1.99Γ—
β€’ CPU generate: 13.8 vs 4.4 β†’ 3.13Γ—
β€’ Metal prompt: 240.7 vs 73.4 β†’ 3.28Γ—

One more quiet fact: the same-size, quality-oriented rival Ternary-Bonsai-27B (7.2 GB) fails to load in upstream llama.cpp at all β€” it needs the PrismML fork. POCKET runs on the tools you already have: LM Studio, Ollama, PocketPal, MLX.

πŸ“– Full story (tech, measurements, recipes): https://huggingface.co/blog/FINAL-Bench/pocket

Models:
πŸ“¦ POCKET-35B-GGUF (PC / server, no GPU): FINAL-Bench/POCKET-35B-GGUF
πŸ‡°πŸ‡· POCKET-KR-GGUF (Android): FINAL-Bench/POCKET-KR-GGUF
🍎 POCKET-KR-MLX (iPhone / Mac): FINAL-Bench/POCKET-KR-MLX
🌍 POCKET-EN-GGUF (English phone / PC): FINAL-Bench/POCKET-EN-GGUF
πŸ–₯️ Live demo (answering on a CPU, no GPU): FINAL-Bench/POCKET-35B-CPU
πŸ“š Collection: FINAL-Bench/pocket-models-6a618ee5d23eafb7e185a5c6
  • 5 replies
Β·
SeaWolf-AIΒ 
posted an update 14 days ago
view post
Post
5142
A small gift for anyone building or studying foundation models.

Most "open" models hand you the weights and stop there. With Aether-7B-5Attn we wanted to hand over the whole thing β€” so you can actually learn from it, reproduce it, and build on it: the data recipe, the training code, every hyperparameter, the complete logs, and the intermediate checkpoints. All Apache-2.0, reproducible byte-for-byte.

What you can do with it:
πŸ” Rebuild it from scratch, or fork the recipe for your own model
πŸ”¬ Study a real heterogeneous-attention MoE β€” 49 layers place 5 attention mechanisms on a 7Γ—7 Latin square, arranged as a clean, attributable ablation
πŸ“ˆ Trace training dynamics across the released checkpoints (110k / 115k / 162k)

It's a modest 6.59B model, and an honest one β€” the limitations (no KV-cache in this build, small scale) are written right in the card. We're not claiming it's special. If any piece of it saves you time or teaches you something, that's exactly what we hoped for. πŸ€—

πŸ“– Full write-up β†’
[blog] Β· https://huggingface.co/blog/FINAL-Bench/opensource-llm
πŸ“¦ 5 Attention Base Β· FINAL-Bench/Aether-7B-5Attn
🎯 5 Attention Instruct · FINAL-Bench/Aether-7B-5Attn-it
πŸš€ 5 Attention Live demo Β· FINAL-Bench/Aether-Sovereign-AI
πŸ“¦ 7 Attention Base Β· https://huggingface.co/FINAL-Bench/Aether-7B-7Attn-base
πŸ“¦ 11 Attention Base Β· FINAL-Bench/Aether-6B-11Attn-base
🧬 Collection · https://huggingface.co/collections/FINAL-Bench/aether-foundation-model

#opensource #LLM #MoE #reproducibility #Apache2
  • 5 replies
Β·
SeaWolf-AIΒ 
posted an update 21 days ago
view post
Post
5404
πŸ”΅ VKUE β€” No GPU? Runs anyway.

"Frontier models need a datacenter GPU" rests on a hidden assumption: that the model reads ALL its parameters every token. Decode is memory-bandwidth bound β€” sweep 34B params/token and an 8 GB card dies at 1–2 tok/s.

So we ran ONE 34.7B reasoning model β€” Ourbox-35B-JGOS, a sparse Mixture-of-Experts β€” as the identical weights across the whole hardware spectrum. All measured:

β€’ B200: 18,057 tok/s (aggregate)
β€’ 1Γ— A10G: 126 tok/s
β€’ 8 GB laptop (RTX 5060): 20 tok/s
β€’ GPU-less CPU: 17 tok/s

Why it works: Ourbox holds 34.7B params but only ~3B are active per token (256 experts, top-8). Since decode is bandwidth-bound, a dense 34B moves ~16.7 GB/token while Ourbox moves ~1.45 GB β€” ~11Γ— less traffic. Put the experts in system RAM, keep attention/router/shared on the GPU, and a 34.7B reasoner runs on an 8 GB laptop β€” or no GPU at all.

Sparsity alone, proven (same laptop, same quant, ~same footprint): Ourbox-35B (A3B) 20.01 tok/s vs Qwen2.5-32B (dense) 5.36 β†’ 3.7Γ— from sparsity alone, ~2Γ— the best dense-32B on any 8 GB machine. Not a toy: GPQA Diamond 86.4% (maj@8).

Try it live (same prompt, GPU vs GPU-less CPU, live tok/s). Honest scope: one machine's measurements; the CPU path proves it RUNS without a GPU, not that it beats one.

πŸ“ Article: https://huggingface.co/blog/FINAL-Bench/vkue
πŸ”΅ GPU vs CPU demo: https://huggingface.co/proxy/final-bench-ourbox-35b-vkue-demo.hf.space/
πŸ”΅ CPU-only demo: https://huggingface.co/proxy/final-bench-ourbox-35b-vkue-cpu.hf.space
πŸ“Š VKUE leaderboard: FINAL-Bench/VKUE
πŸ€— Model: FINAL-Bench/Ourbox-35B-JGOS-GGUF
⚑ VKAE (speed): VIDraft/vkae

VKUE is the "runs anywhere" side of our serving line; VKAE the "fast on datacenter GPUs" side. VKAE is fast; VKUE is everywhere.
  • 7 replies
Β·
SeaWolf-AIΒ 
posted an update 28 days ago
view post
Post
5193
πŸ”“ We ran genuine quantum key-recovery on 'real IBM quantum hardware' β€” and pushed the frontier well past the largest hardware demos we're aware of (which sat at N=4).

Using Simon's algorithm on ibm_kingston, we recovered the secret key of two symmetric-cipher structures:
β€’ Even–Mansour β€” N=5 β†’ N=10
β€’ 3-round Feistel (DES-family) β€” block 6 β†’ 8

Each verified against an 'independent control key', using error mitigation only (no QEC).

🧭 Honest scope: this is not a quantum speedup (the effective difficulty tracks the classical birthday bound ~2^{n/2}), not a break of real AES/RSA, and not 16-round DES (ours is 3-round). The recovery method is reserved for a forthcoming paper; formal record status is pending peer review.

πŸ“„ Write-up: https://huggingface.co/blog/FINAL-Bench/quantum
πŸ•ΉοΈ Try it live in your browser: https://huggingface.co/proxy/vidraft-quantumos.hf.space/crypto
πŸ† Leaderboard: FINAL-Bench/quantum-bench-leaderboard

#quantum #cryptography #quantumcomputing
SeaWolf-AIΒ 
posted an update 30 days ago
view post
Post
3009
πŸš€ Adding a GPU without building one

AI is usually framed as "how smart is the model / how many GPUs did you buy." The real bottleneck is elsewhere β€” how efficiently you use the GPUs you already have.

Training happens once; inference runs the entire time users use your product. So a service's economics come down to cost per token. Inference acceleration uses software to pull several times more out of the same GPU β€” the effect of plugging in one more "virtual GPU."

VIDRAFT's VKAE, measured (B200, same-harness, no quality loss):

Qwen3.5-35B-A3B (MoE): 25.7 β†’ 601 tok/s (23.4Γ—)
Darwin-36B-Opus (in-house MoE): 25.0 β†’ 280.8 (11.2Γ—)
10,000+ tok/s peak aggregate under concurrency
The key: it's reproducible β€” model + serving shipped as one container.

docker pull vidraft/qwen35-vkae:601
Don't take our word for it β€” run it yourself. The mechanism will be released as a paper.

πŸ† Leaderboard & demo πŸ‘‰ VIDraft/vkae
Articles πŸ‘‰ https://huggingface.co/blog/FINAL-Bench/vkae-leaderboard
  • 3 replies
Β·
SeaWolf-AIΒ 
posted an update about 1 month ago
view post
Post
5227
🐯 Chitos β€” The Security Scanner That Actually Proves It

Most security scanners hand you a suspect list and walk away. That gap between detection and proof is where attackers live β€” and it's exactly the gap that Chitos was built to close.

Chitos is the successor to Mythos, a static analyzer built for quick code health checks. Mythos was good at pattern matching β€” spotting dangerous sinks, mapping CWEs, producing readable reports. But static analysis has a structural ceiling. A rule that sees eval(user_input) can tell you that looks dangerous. It cannot tell you whether the input is reachable, whether sanitization three layers up covers this path, or whether there's a live exploit chain for your exact framework version. Chitos was built to answer those questions.

πŸ” Phase 1 applies 50 language-agnostic rules across Python, JavaScript, Go, Java, C/C++, Rust, PHP, YAML and more β€” covering injection sinks, deserialization gadgets, credential leakage, broken crypto, and prototype pollution. Every candidate is re-verified before reaching the report. Findings that can't be substantiated are excluded, not handed to you as noise.

πŸ”¬ Phase 2 dispatches an autonomous web-search agent to hunt live CVE databases, exploit advisories, and public PoC repositories. It formulates hypotheses, verifies them, and synthesizes a structured threat narrative. This phase needs a user-supplied Claude API key β€” Phases 1 and 3 run entirely free.

🎯 Phase 3 is where Chitos diverges from everything else. Against targets you own or are authorized to test, it fires real payloads β€” XSS, SQLi, path traversal, command injection β€” mutates on block, captures hard evidence, and connects every proven finding into a kill-chain showing which vulnerabilities to remediate first.

No installation. No account. No code sent to third-party APIs.

Article: https://huggingface.co/blog/FINAL-Bench/chitos

Try it now πŸ‘‰ https://chitos.vidraft.net
  • 10 replies
Β·
SeaWolf-AIΒ 
posted an update about 2 months ago
view post
Post
4930
Darwin V9 β€” GPQA Diamond 90.9%, #1 on the leaderboard, with pure greedy decoding
Darwin-398B-JGOS reaches 90.9% (180/198) on GPQA Diamond, the PhD-level scientific reasoning benchmark, ranking #1 on the Hugging Face GPQA Diamond leaderboard. No self-consistency, no test-time compute scaling β€” this was achieved with a single greedy decode (temperature 0, single sample, max 16,384 tokens). The full eval config is published in the model card, so anyone can reproduce it. Raw reasoning, no score inflation.
The result comes from Darwin V9, a patented evolutionary model-development platform. Its core idea: it never trains a model from scratch.
Why Darwin V9 beats training from scratch

Cost & speed: no trillion-token pretraining run, no months of compute β€” a purpose-built, high-performance model is produced in a fraction of the time.
Reuse of proven intelligence: instead of re-learning every capability from a blank slate, it selects and combines only the strengths of already-trained, already-validated models, so results are stable and predictable.
Surgical transplantation: it identifies which neural region of which model holds which capability β€” at the FFN (Feed Forward Network) layer level β€” and grafts in only the segments that contribute to the target skill.

How it works: a large model (Qwen 3.5 397B) serves as the mother model (the substrate); several father models specialized in reasoning, coding, and language are analyzed layer-by-layer across their FFN regions; the segments that contribute to the target performance are extracted and transplanted into the mother model to produce a new child model. The result is a ~400B MoE that activates only ~17B parameters per token at inference β€” large-model capacity with efficient inference.
If training from scratch means rebuilding everything from a blank page, Darwin V9 means precisely recombining intelligence that has already been proven. GPQA Diamond #1 is the proof.
Model: FINAL-Bench/Darwin-398B-JGOS
SeaWolf-AIΒ 
posted an update about 2 months ago
view post
Post
6803
πŸš€ Introducing FINAL-Bench Quantum β€” an open, neutral benchmark that finally puts quantum-computing methods on one fair yardstick.

Quantum results are notoriously hard to compare. The same "logical error rate" or "query fidelity" means very different things depending on the code, noise model, hardware, and shot count. FINAL-Bench Quantum fixes that: five events judged under identical, published protocols, where every number is labeled as either measured here or quoted from a source.

Five events: β‘  QEC Decoder β‘‘ Optimization (Max-Cut) β‘’ VQE β‘£ QRAM β‘€ Quantum Simulation

The rules are simple and strict:
βœ… Track A (measured here, with 95% confidence intervals) is kept separate from Track B (quoted from papers, not directly comparable).
πŸ”¬ Simulation and real hardware are clearly distinguished, and no quantum-advantage claims are made.
🌍 Methods from Google, IBM, NVIDIA, USTC, Riverlane and more sit side by side, with origin flags and author credits.
πŸ“€ Anyone can submit their own method via the Submit tab for review and listing.

Already on the board: real IBM Heron r2 measurements (repetition-code distance boundary, 29–175Γ— error reduction from d3 to d5), a real-chip QRAM query fidelity of 0.92, and Hβ‚‚ VQE at chemical accuracy β€” always labeled honestly as simulation vs hardware.

A leaderboard is only useful if you can trust it, so neutrality is the whole point: strong competitors stay in even when they beat the host, sources are quoted faithfully, and a simulation is never rounded up into a hardware claim.

Leaderboard: FINAL-Bench/quantum-bench-leaderboard
Article: https://huggingface.co/blog/FINAL-Bench/quantum-leaderboard

#quantum #QEC #QuantumComputing #benchmark
  • 2 replies
Β·
SeaWolf-AIΒ 
posted an update 2 months ago
view post
Post
4271
Darwin-60B-DUO: Two SOTAs, One Endpoint β€” 88.38% on GPQA Diamond πŸš€

We're excited to release Darwin-60B-DUO, the Darwin family's first DUO model. Take two domain-verified specialists, hide them behind a single OpenAI-compatible endpoint, and let a router decide which one (or both) answers. You see one model, one API β€” but get the best of both.

The number that matters: on the full 198-question GPQA Diamond, Darwin-60B-DUO hits 88.38%. The constituents alone land at 69.70% (Darwin-28B-REASON) and 77.27% (AWAXIS-Think-31B); a naive cascade only reaches 83.84%. The DUO clears them all. Two small specialists, intelligently routed, beat one big generalist on cost and quality. Both are independently verified β€” Darwin-28B-REASON is #3 on the HF GPQA Diamond leaderboard, AWAXIS-Think-31B is #1 on Korea's national K-AI Leaderboard (MSIT).

The brains is a Hybrid-A router picking one of five strategies on the fly. Korean β†’ AWAXIS, English/STEM β†’ Darwin (single-backend, ~70% of traffic at 1Γ— cost). When a Korean answer needs rigorous English reasoning, split_refine fires β€” Darwin drafts, AWAXIS polishes; MCQ/short-answer runs both with self-consistency + cross-verify. Net effective cost: only ~1.3Γ— a single 30B model.

The part the community will care about: the gateway is model-agnostic and Apache-2.0. Point it at any two OpenAI-compatible backends and you've got a DUO in minutes β€” teach router.py when to use which, and parallel calls, response merging, and routing transparency via _duo_route are handled for you. Fork it and tell us what you built.

Painless deploy: docker compose up for both vLLM backends + gateway; FP8 ~30GB colocates on a single B200/H100. One git clone (~120GB). Text-only for now, streaming in v1.1.
Two SOTAs, one endpoint. Come build your own on the Community tab.

πŸ‘‡
πŸ”— FINAL-Bench/Darwin-60B-DUO
SeaWolf-AIΒ 
posted an update 3 months ago
view post
Post
5438
🧬 Darwin Family: Zero Gradient Steps, GPQA Diamond 88.89%

How far can we push LLM reasoning *without* training?

Our team at VIDRAFT submitted this paper to Daily Papers yesterday, and it's
currently #3. Huge thanks to everyone who upvoted β€” sharing the core ideas below.

πŸ”— Paper: Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning (2605.14386)
πŸ”— arXiv: https://arxiv.org/abs/2605.14386
πŸ”— Model: FINAL-Bench/Darwin-28B-REASON
πŸ”— Model: FINAL-Bench/Darwin-28B-Opus

---

TL;DR

Darwin Family is a training-free evolutionary merging framework.
By recombining the weight spaces of existing LLM checkpoints β€” with zero
gradient-based training β€” it reaches frontier-level reasoning.

- πŸ† Darwin-28B-Opus: GPQA Diamond 88.89%
- πŸ’Έ Zero gradient steps β€” not a single B200 or H200 hour needed
- 🧬 Consistent gains across 4B β†’ 35B scale
- πŸ”€ Cross-architecture breeding between Transformer and Mamba families
- πŸ” Stable recursive multi-generation evolution

#Three Core Mechanisms

β‘  14-dim Adaptive Merge Genome β€” fine-grained recombination at both
component level (Attention / FFN / MLP / LayerNorm / Embedding) and block
level, expanding the prior evolutionary-merge search space.

β‘‘ MRI-Trust Fusion β€” we diagnose each layer's reasoning contribution
via an **MRI (Model Reasoning Importance)** signal and fuse it with
evolutionary search through a **learnable trust parameter**. Trust the
diagnostic too much and search collapses; ignore it and search becomes
inefficient β€” Darwin learns the balance from data.

β‘’ Architecture Mapper β€” weight-space breeding across heterogeneous
families. Attention Γ— SSM crossover actually works.

Why It Matters
> Diagnose latent capabilities already encoded in open checkpoints,
> and recombine them β€” no gradients required.

Replies and critiques welcome πŸ™Œ
  • 3 replies
Β·
SeaWolf-AIΒ 
posted an update 3 months ago
view post
Post
5103
🌌 Introducing Model Galaxy β€” a Living, Multimodal Fork of the HF Model Atlas

πŸ‘‰ Try it: FINAL-Bench/model-galaxy

This Space is a fork of the brilliant Eliahu/Model-Atlas, the official demo of "Charting and Navigating Hugging Face's Model Atlas" (Horwitz et al., arXiv 2503.10633). Their pre-computed HF model graph is the foundation of every node and edge you see, and we are deeply grateful for its open release.

The original atlas is a static snapshot of early 2025. Model Galaxy turns it into a living, multimodal map. We injected the 2026 trending originals that did not exist when the atlas was frozen β€” DeepSeek-V4, Hy3-preview, GLM-5.1, Kimi-K2, gpt-oss, Nemotron-3 Super / Nano / Omni, Hermes-4.3, Qwen3-Coder-Next, Llama-3.3, Granite-4.1, plus the latest multimodal releases (FLUX.2, ERNIE-Image, HunyuanImage / Video, LTX-2.3, Wan2.2, Kokoro-82M, VoxCPM2, Voxtral-TTS, whisper-v3-turbo, Gemma-4, Qwen3-Omni, Phi-4-mm) β€” each with proper base_model lineage edges.

We also added the complete VIDRAFT Darwin family ontology: 120 nodes covering Darwin Core, AETHER, every brand variant (Rogue, AWAXIS, TenOS, Warecube), NOESIS-Darwin multimodal extensions, and 40+ community quantizations β€” the most complete Darwin lineage view anywhere.

The name "Galaxy" is now literal: our three injected clusters are re-laid out as logarithmic spiral galaxies, with bigger models near the bright cores and quantizations scattering to the outer arms β€” just like real star mass distribution. A top-right toggle switches between Galaxy mode (deep-space gradient with 220 animated stars) and Atlas mode (clean white panels for reports). A 15-second progress bar narrates the render, and per-modality / per-company colors make every cluster legible at a glance.

Final scale: 22,480 nodes in the default Modalities atlas, 137,324 in the Large NLP atlas, and a 277-node compact Darwin + Trending view for instant exploration. Feedback and PRs welcome.
SeaWolf-AIΒ 
posted an update 3 months ago
view post
Post
8798
🧬 Introducing Darwin-9B-NEG β€” the first model with Native Entropy Gating (NEG)

πŸ”— Try it now: FINAL-Bench/Darwin-9B-NEG
πŸ”— Q4 bit : FINAL-Bench/Darwin-9B-MFP4

We're thrilled to release Darwin-9B-NEG, a 9B-parameter reasoning model
that embeds an architecturally-internalised sense of self-confidence directly
into the transformer β€” our proprietary Native Entropy Gating (NEG) technology.

πŸ“Š GPQA Diamond (198 PhD-level questions):

β–Έ Baseline Darwin-9B (no NEG) β†’ 51.01 %
β–Έ Pure NEG (greedy Β· 1Γ— cost) β†’ 63.64 % πŸ”₯ +12.63 %p
β–Έ + Permutation (4Γ— cost) β†’ 76.26 %
β–Έ + Ensemble Refinement (~20Γ—) β†’ 84.34 % πŸ†

With only 9 billion parameters and 1Γ— inference cost, Pure NEG jumps
+12.63 %p over the same model without NEG. Going all-in with ensemble
refinement pushes it to 84.34 % β€” surpassing the published Qwen3.5-9B
leaderboard score (81.7 %) by +2.64 %p.

πŸ”¬ What makes NEG different from Multi-Turn Iteration (MTI)?

Classical MTI needs 3-8Γ— extra inference passes. NEG instead lives
INSIDE the single decoding loop. Two tiny modules ride with the
transformer: NEG-Head predicts per-token entropy from the last hidden
state, and NEG-Gate conditionally restricts the top-k choice when
confidence is low. The gate activates in only 4.36 % of tokens β€”
essentially free at inference time.

✨ Key differentiators
β€’ Architecturally internalised β€” model file *is* the feature
β€’ 1Γ— inference cost (vs. 3-8Γ— for MTI)
β€’ Drop-in with vLLM / SGLang / TGI / transformers β€” no extra engine
β€’ +12.63 %p reasoning at zero latency overhead
β€’ Single-file deployment, Apache 2.0 licensed

🧬 Lineage
Qwen/Qwen3.5-9B β†’ Darwin-9B-Opus (V7 evolutionary merge) β†’ Darwin-9B-NEG (V8 + NEG training)

#Darwin #NEG #NativeEntropyGating #GPQA #Reasoning #LLM #OpenSource #Apache2
SeaWolf-AIΒ 
posted an update 4 months ago
view post
Post
4476
Darwin-TTS: 3% of an LLM's Brain Makes TTS Speak with Emotion β€” Zero Training

We blended 3% of Qwen3-1.7B (LLM) FFN weights into Qwen3-TTS-1.7B's talker module. The result: emotionally enhanced speech synthesis β€” with zero training, zero data, and zero GPU hours.

Try the Demo: FINAL-Bench/Darwin-TTS-1.7B-Cross

Model Weights: FINAL-Bench/Darwin-TTS-1.7B-Cross

Full Research Article: https://huggingface.co/blog/FINAL-Bench/darwin-tts

Qwen3-1.7B (LLM) and Qwen3-TTS-1.7B's talker share 100% identical architecture β€” same hidden_size (2048), same layers (28), same heads (16). This enabled pure 1:1 weight blending across 84 FFN tensors with a single lerp operation. At 3% blend, emotion appears. At 5%, emotion intensifies. At 10%, the model breaks β€” producing 655-second outputs for a 3-second sentence, because the LLM's "keep generating" pattern overwhelms the TTS stop signal.

To our knowledge, this is the first training-free cross-modal weight transfer between an LLM and a TTS model. Prior work either requires adapter training (SmolTolk, 2025), fine-tuning (CSLM, 2025), or massive end-to-end compute (GPT-4o). Darwin-TTS achieves cross-modal capability transfer in under 2 minutes on CPU.

The key insight: TTS models with LLM backbones already "think" in language. We're just restoring 3% of the original LLM's language understanding patterns β€” particularly those related to emotional semantics and prosody planning. The code is three lines: load the model, load the LLM FFN, call p.lerp_(llm_weight, 0.03).

creators of the Darwin Evolutionary Merge Framework.
Darwin LLM V7 achieved GPQA Diamond 86.9% (HF Benchmark #3)
through CMA-ES optimized FFN crossbreeding. Darwin-TTS extends this principle from LLM-to-LLM merging into cross-modal LLM-to-TTS transfer. Apache 2.0.
SeaWolf-AIΒ 
posted an update 4 months ago
view post
Post
5977
🧬 Darwin-27B-Opus: 86.9% on GPQA Diamond β€” World #5, Zero Training
We are excited to share Darwin-27B-Opus, a 27B model that achieved 86.9% on GPQA Diamond β€” ranking #5 globally on the HuggingFace leaderboard β€” without a single gradient update.

How? Darwin breeds pretrained models through evolutionary FFN crossbreeding. The father (Qwen3.5-27B) provides the reasoning architecture; the mother (Claude 4.6 Opus Reasoning Distilled) contributes structured chain-of-thought knowledge. CMA-ES automatically discovers optimal per-layer blending ratios β€” no human tuning required.

The result surpasses the original Qwen3.5-27B (85.5%), GLM-5.1 (744B, 86.2%), and Qwen3.5-122B (86.6%). A 27B model outperforming 744B β€” with zero training, zero data, one GPU, ~2 hours.

We also confirmed hybrid vigor on Korean benchmarks: Darwin-27B-KR (2nd generation offspring) surpassed both parents on CLIcK, winning 7 out of 11 categories. The evolutionary optimizer independently assigned 93% of FFN from the Korean-specialized mother while preserving 93% of attention from the reasoning-specialized father β€” autonomously validating our core principle: FFN carries knowledge, Attention carries reasoning.

πŸ“Š Public release: 10 days β†’ 300+ community derivatives, 120K+ downloads.

πŸ”— Links:
Darwin-27B-Opus: FINAL-Bench/Darwin-27B-Opus
article: https://huggingface.co/blog/FINAL-Bench/darwin-gpqa
Darwin Family Collection: https://huggingface.co/collections/FINAL-Bench/darwin-family

If foundation models are raw ore, Darwin is the forge. We are just getting started. πŸ”₯
SeaWolf-AIΒ 
posted an update 4 months ago
view post
Post
3065
Why This Matters β€” David Defeats Goliath

MODEL: FINAL-Bench/Darwin-4B-David
SPACE: FINAL-Bench/Darwin-4B-david

We're releasing Darwin-4B-David, the first second-generation model in the Darwin Opus family. By evolving an already-evolved model, it achieves 85.0% on GPQA Diamond β€” surpassing its 58.6% original ancestor and even gemma-4-31B (84.3%) β€” with just 4.5B parameters.

Second-Generation Evolution
Most merges start from a base model and produce a single offspring. Darwin-4B-David breaks this pattern. The Father (Darwin-4B-Opus) was already evolved from gemma-4-E4B-it with Claude Opus reasoning distillation β€” a Gen-1 model. The Mother (DavidAU's DECKARD-Expresso-Universe) brings Unsloth deep tuning across 5 in-house datasets with thinking mode by default. Crossbreeding these two produced the first Gen-2 Darwin model.

Darwin V6's Model MRI scanned both parents across all 42 layers, assigning independent optimal ratios per layer. The Mother's creativity and Korean language hotspot (Layer 22-25, weight 0.95) was maximally absorbed, while the Father's reasoning core (Layer 30-40, weight 0.48) was preserved. This is "Merge = Evolve" applied recursively β€” evolution of evolution.

Benchmarks
Darwin-4B-David scores 85.0% on GPQA Diamond (+26.4%p over original 58.6%), evaluated generatively with maj@8 (8 generations per question, majority vote), Epoch AI prompt format, thinking mode enabled, 50 sampled questions. On ARC-Challenge (25-shot, loglikelihood), both score 64.93% β€” expected, as loglikelihood doesn't capture thinking-mode reasoning differences.

Why This Matters
gemma-4-31B (30.7B) scores 84.3%. Darwin-4B-David surpasses it at 1/7th the size β€” no training, no RL, just 45 minutes of MRI-guided DARE-TIES on one H100. The name "David" honors Mother creator DavidAU and evokes David vs. Goliath.
SeaWolf-AIΒ 
posted an update 4 months ago
view post
Post
5591
🧬 Darwin V6: Diagnostic-Guided Evolutionary Model Merging

We are releasing Darwin-31B-Opus β€” a reasoning-enhanced model merging Google's Gemma-4-31B-it and TeichAI's Claude Opus Distill using the Darwin V6 engine.

Model: FINAL-Bench/Darwin-31B-Opus
Demo: FINAL-Bench/Darwin-31B-Opus

πŸ”¬ What Darwin V6 Does

Conventional merging tools (mergekit, etc.) apply a single ratio to all tensors. Set ratio=0.5 and all 1,188 tensors blend identically, with no distinction between which tensors matter for reasoning versus coding.

Darwin V6 diagnoses both parents at the tensor level before merging. It measures Shannon entropy, standard deviation, and L2 norm for every tensor, then passes 5 diagnostic probes (REASONING, CODE, MATH, KNOWLEDGE, LANGUAGE) through the model to determine layer-wise functional importance. Each of the 1,188 tensors receives an independent optimal ratio.

combined = static(entropy/std/norm) x 0.4 + probe(cosine_distance) x 0.6
final_ratio = mri_ratio x mri_trust + genome_ratio x (1 - mri_trust)

When one parent is overwhelmingly superior for a tensor (ratio < 0.15 or > 0.85), Darwin transplants it directly without interpolation. The mri_trust parameter itself is optimized by CMA-ES evolutionary search, so optimal transplant intensity is determined automatically. After merging, a Health Check compares the child against both parents layer-by-layer to detect interference or function loss.

🧬 Parent Models
Father: google/gemma-4-31B-it
Mother: TeichAI/gemma-4-31B-it-Claude-Opus-Distill

🧬 Results
Compared under identical conditions (same 50 questions, same seed, greedy, thinking mode):
Father: 60.0% (30/50)
Darwin-31B-Opus: 66.0% (33/50) β€” +10% relative improvement
ARC-Challenge: 82.89% (loglikelihood, zero-shot, 200 questions)
Optimal genome found by evolution:
ffn_ratio=0.93 β€” FFN layers strongly favor Mother (Claude Opus Distill)
block_5 (L50-L59)=0.86 and more...
  • 11 replies
Β·
SeaWolf-AIΒ 
posted an update 4 months ago
view post
Post
3218
πŸ’Ž Gemma 4 Playground β€” Dual Model Demo on ZeroGPU

We just launched a Gemma 4 Playground that lets you chat with Google DeepMind's latest open models β€” directly on Hugging Face Spaces with ZeroGPU.

FINAL-Bench/Gemma-4-Multi

πŸ‘‰ Try it now: FINAL-Bench/Gemma-4-Multi
Two Models, One Space
Switch between both Gemma 4 variants in a single interface:

⚑ Gemma 4 26B-A4B β€” MoE with 128 experts, only 3.8B active params. 95% of the 31B's quality at ~8x faster inference. AIME 88.3%, GPQA 82.3%.
πŸ† Gemma 4 31B β€” Dense 30.7B. Best quality among Gemma 4 family. AIME 89.2%, GPQA 84.3%, Codeforces 2150. Arena open-model top 3.

Features

Vision β€” Upload images for analysis, OCR, chart reading, document parsing
Thinking Mode β€” Toggle chain-of-thought reasoning with Gemma 4's native <|channel> thinking tokens
System Prompts β€” 6 presets (General, Code, Math, Creative, Translate, Research) or write your own
Streaming β€” Real-time token-by-token response via ZeroGPU
Apache 2.0 β€” Fully open, no restrictions

Technical Details
Built with the dev build of transformers (5.5.0.dev0) for full Gemma 4 support including multimodal apply_chat_template, variable-resolution image processing, and native thinking mode. Runs on HF ZeroGPU with @spaces .GPU β€” no dedicated GPU needed.
Both models support 256K context window and 140+ languages out of the box.

Links

- πŸ€— Space: [FINAL-Bench/Gemma-4-Multi]( FINAL-Bench/Gemma-4-Multi)
- πŸ“„ Gemma 4 26B-A4B: [google/gemma-4-26B-A4B-it]( google/gemma-4-26B-A4B-it)
- πŸ“„ Gemma 4 31B: [google/gemma-4-31B-it]( google/gemma-4-31B-it)
- πŸ”¬ DeepMind Blog: [Gemma 4 Launch](https://deepmind.google/blog/gemma-4-byte-for-byte-the-most-capable-open-models/)
  • 2 replies
Β·
SeaWolf-AIΒ 
posted an update 4 months ago
view post
Post
2197
🧬 Darwin-35B-A3B-Opus β€” The Child That Surpassed Both Parents

What if a merged model could beat both its parents? We proved it can.
Darwin-35B-A3B-Opus is a 35B MoE model (3B active) built with our Darwin V5 engine β€” the first evolution system that CT-scans parent models before merging them.
πŸ€— Model: FINAL-Bench/Darwin-35B-A3B-Opus

The result speaks for itself: GPQA Diamond 90.0%, versus Father (Qwen3.5-35B-A3B) at 84.2% and Mother (Claude 4.6 Opus Distilled) at 85.0%. That's +6.9% over Father and +5.9% over Mother. Not a tradeoff β€” a genuine leap. Meanwhile, MMMLU sits at 85.0% (Father: 85.2%), multimodal is fully intact, and all 201 languages are preserved.

How? Model MRI changed everything. Traditional merging is guesswork. Darwin V4 added evolution. Darwin V5 added X-ray vision. Model MRI scans each parent layer by layer and discovers: Mother's L34–L38 is the reasoning engine (peak cosine distance), 50–65% of Mother's experts are dead (killed by text-only distillation), and Father is a healthy generalist with every expert alive. The prescription: transplant Mother's reasoning brain at L38 (90% weight), replace her dead experts with Father's living ones, and let Father's router handle the output layer. Reasoning went up. Versatility stayed intact. No tradeoff β€” just evolution.

35B total, 3B active (MoE) Β· GPQA Diamond 90.0% Β· MMMLU 85.0% (201 languages) Β· Multimodal Image & Video Β· 262K native context Β· 147.8 tok/s on H100 Β· Runs on a single RTX 4090 (Q4) Β· Apache 2.0
Darwin V5's full algorithm and technical details will be released alongside an upcoming paper.

πŸš€ Live Demo: FINAL-Bench/Darwin-35B-A3B-Opus

πŸ† FINAL Bench Leaderboard: FINAL-Bench/Leaderboard

πŸ“Š ALL Bench Leaderboard: FINAL-Bench/all-bench-leaderboard

Built by VIDRAFT Β· Supported by the Korean Government GPU Support Program
  • 8 replies
Β·