Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🔄
In a Training Loop
49.8
TFLOPS
Dean Byrne
PRO
Quazim0t0
23
7
90
Follow
RebelBroUncensored's profile picture
0001AMA's profile picture
C50BARZ's profile picture
230 followers
·
2,497 following
https://huggingface.co/DaisyChainAI
quzi93
dean-byrne-02a28b191
AI & ML interests
DaisyChainAI🌼 / SmallLM's / San Francisco / Open Source
Recent Activity
liked
a model
1 day ago
opencerebral/Boris-1.7-D60M-n30M
reacted
to
FlameF0X
's
post
with 🔥
1 day ago
Hello HuggingFace! I would like to share a small preview of a language model that I have been experiencing for a while. In the first attached image there is a small sample of the Myosotis-1, an attempt to make a small 100m parameter flagship model that is built on my bizarre architecture that is somewhat similar to an S4/S5 model with WKV added. I call it FWKV (Feed-Forward WKV). The model is currently still in training because of the nature of RNN-like models. It can also be seen that the model has insane prompt processing and token generation speed (evaluation done on a 2x Titan XP); even for its small size, some similar Transformer models do struggle to get the same results without custom kernels (some Transformer models can achieve this level of throughput on cheap hardware). In the second image, you can see checkpoint 20k of the model in its next token prediction state (this means it can't chat), ranking in the top 100 on https://huggingface.co/spaces/AxiomicLabs/Open_SLM_Leaderboard (the results have not been submitted since the model is not done training). - Why not just use Transformers? Have you seen any pure non-Transformers SLMs besides RWKV and Mamba? - Should you expect this project to become the next LFM or another very fast language model thing on some Raspberry Pi? No, the model is still an experiment; it's very sensible and prone to collapse (by the time of this post, it can be seen in the 1st image). - Should you use it? Maybe not yet; the architecture itself is still very "naive"—that's how I could call it at its current level. If you just want to play with it and see what you could do or how fast the model is on your hardware, then you can do it. Once the training is finished and I feel satisfied with the model next token prediction (the base model) and "assistants" (the instruction-tuned model) capabilities, I will make open weights at https://huggingface.co/FWKV with full support of the 🤗 Transformers.
reacted
to
FlameF0X
's
post
with 👍
1 day ago
Hello HuggingFace! I would like to share a small preview of a language model that I have been experiencing for a while. In the first attached image there is a small sample of the Myosotis-1, an attempt to make a small 100m parameter flagship model that is built on my bizarre architecture that is somewhat similar to an S4/S5 model with WKV added. I call it FWKV (Feed-Forward WKV). The model is currently still in training because of the nature of RNN-like models. It can also be seen that the model has insane prompt processing and token generation speed (evaluation done on a 2x Titan XP); even for its small size, some similar Transformer models do struggle to get the same results without custom kernels (some Transformer models can achieve this level of throughput on cheap hardware). In the second image, you can see checkpoint 20k of the model in its next token prediction state (this means it can't chat), ranking in the top 100 on https://huggingface.co/spaces/AxiomicLabs/Open_SLM_Leaderboard (the results have not been submitted since the model is not done training). - Why not just use Transformers? Have you seen any pure non-Transformers SLMs besides RWKV and Mamba? - Should you expect this project to become the next LFM or another very fast language model thing on some Raspberry Pi? No, the model is still an experiment; it's very sensible and prone to collapse (by the time of this post, it can be seen in the 1st image). - Should you use it? Maybe not yet; the architecture itself is still very "naive"—that's how I could call it at its current level. If you just want to play with it and see what you could do or how fast the model is on your hardware, then you can do it. Once the training is finished and I feel satisfied with the model next token prediction (the base model) and "assistants" (the instruction-tuned model) capabilities, I will make open weights at https://huggingface.co/FWKV with full support of the 🤗 Transformers.
View all activity
Organizations
Quazim0t0
's models
52
Sort: Recently updated
Quazim0t0/Byrne-100M-Ultra-MC
Text Generation
•
0.1B
•
Updated
3 days ago
•
294
•
5
Quazim0t0/Byrne-15M-Looped
Text Generation
•
Updated
9 days ago
•
1
Quazim0t0/Anti-DEGR-Toolkit
Text Generation
•
Updated
9 days ago
Quazim0t0/CS2Dream
Updated
27 days ago
•
1
Quazim0t0/Positronic-144M-Anti-DEG
Text Generation
•
Updated
Jul 30
Quazim0t0/Escarda-86M-Identity-Anti-DEG
Text Generation
•
97.3M
•
Updated
Jul 30
•
192
•
1
Quazim0t0/Byrne-86M-Base-JL-Anti-DEG
Text Generation
•
96.9M
•
Updated
Jul 21
•
32
Quazim0t0/Escarda-86M-Base-JL-Anti-DEG
Text Generation
•
97.3M
•
Updated
Jul 21
•
26
Quazim0t0/Wheeler-DeWitt-62M
Text Generation
•
Updated
Jul 21
•
1
Quazim0t0/BE-ImgGen-Portrait-187M
Text-to-Image
•
Updated
Jul 20
Quazim0t0/Spikewhale-SNN-Brain2Qwerty
Text Generation
•
Updated
Jul 15
Quazim0t0/Byrne-DFINE-N
Object Detection
•
Updated
Jul 12
•
36
Quazim0t0/Onnx-Library
Text Generation
•
Updated
Jul 12
Quazim0t0/neural-photonic
Updated
Jul 11
•
2
Quazim0t0/Escarda-86M-Base-JL
Text Generation
•
97.3M
•
Updated
Jul 11
•
101
Quazim0t0/Byrne-86M-Base-JL
Text Generation
•
96.9M
•
Updated
Jul 11
•
73
Quazim0t0/neural-cd-preserve
Updated
Jul 11
•
1
Quazim0t0/neural-storage
Updated
Jul 11
•
2
Quazim0t0/neural-ddr
Updated
Jul 11
•
1
Quazim0t0/Chimera-64M
Text Generation
•
Updated
Jul 10
•
2
Quazim0t0/Mycel-LM-79M
Text Generation
•
Updated
Jul 10
•
5
Quazim0t0/SpikeWhale-SNN-216M
Text Generation
•
Updated
Jul 10
•
5
Quazim0t0/Positronic-144M
Text Generation
•
Updated
Jul 10
•
6
Quazim0t0/physgait-weights
Reinforcement Learning
•
Updated
Jul 9
Quazim0t0/Byrne-VE
Image Feature Extraction
•
Updated
Jul 9
Quazim0t0/Byrne-VLM-131M
Image-to-Text
•
Updated
Jul 9
Quazim0t0/Byrne-Docling-131M
Image-to-Text
•
Updated
Jul 9
•
2
Quazim0t0/Escarda-Docling-126M
Image-to-Text
•
Updated
Jul 9
•
1
Quazim0t0/Escarda-86M-Identity
Text Generation
•
97.3M
•
Updated
Jul 8
•
105
•
1
Quazim0t0/Byrne-TriAtn-86M
Text Generation
•
96.9M
•
Updated
Jul 8
•
23
•
1
Previous
1
2
Next