Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🔄
In a Training Loop
aquilesfd
aquilesfd
50
22
307
Follow
Datdanboi25's profile picture
appvoid's profile picture
Tonic's profile picture
39 followers
·
59 following
https://aquilesfd.vercel.app
aquilesfd
aquilesfd
aquiles-fernandes-53b7413a5
AI & ML interests
research
Recent Activity
replied
to
Banaxi-Tech
's
post
about 9 hours ago
We're announcing our BananaMind 2.1 model series! The models will include: - BananaMind 2.1 Nano: 10M parameters with 60B tokens. - BananaMind 2.1 Lite: 25M parameters with 40B tokens. - BananaMind 2.1 Flash: 50M parameters with 55B tokens. - BananaMind 2.1 Pro: 135M-145M parameters (still deciding) with 100B tokens. These model will use a multi tower architecture (like https://huggingface.co/BananaMind/BananaMind-2.1-Unified) with some more architectural changes. BananaMind 2.1 Pro will probrably use 2 no output towers, instead of one! We're currently training some experimental models based on this architecture to see its scaling! Follow us: https://huggingface.co/BananaMind @Banaxi-Tech @vovaRL @DedeProGames https://huggingface.co/bananamind-research-community
reacted
to
Banaxi-Tech
's
post
with 🔥
about 10 hours ago
We're announcing our BananaMind 2.1 model series! The models will include: - BananaMind 2.1 Nano: 10M parameters with 60B tokens. - BananaMind 2.1 Lite: 25M parameters with 40B tokens. - BananaMind 2.1 Flash: 50M parameters with 55B tokens. - BananaMind 2.1 Pro: 135M-145M parameters (still deciding) with 100B tokens. These model will use a multi tower architecture (like https://huggingface.co/BananaMind/BananaMind-2.1-Unified) with some more architectural changes. BananaMind 2.1 Pro will probrably use 2 no output towers, instead of one! We're currently training some experimental models based on this architecture to see its scaling! Follow us: https://huggingface.co/BananaMind @Banaxi-Tech @vovaRL @DedeProGames https://huggingface.co/bananamind-research-community
liked
a model
1 day ago
SupraLabs/Supra2-Medium-Base
View all activity
Organizations
aquilesfd
's models
8
Sort:Â Recently updated
aquilesfd/neo-3-1B-A90M-Instruct
Text Generation
•
Updated
Apr 17
•
8
aquilesfd/neo-3-3B-A400M-Thinking
Text Generation
•
Updated
Feb 13
•
3
aquilesfd/neo-3-1B-A90M-Base
Text Generation
•
1.0B
•
Updated
Feb 13
•
189
•
4
aquilesfd/neo-3-3B-A400M-Base
Text Generation
•
3B
•
Updated
Feb 13
•
104
•
4
aquilesfd/neo-2-345M-C2
Text Generation
•
0.4B
•
Updated
Dec 30, 2025
•
65
•
1
aquilesfd/neo-2-345M-C1
Text Generation
•
0.4B
•
Updated
Dec 30, 2025
•
70
•
1
aquilesfd/neo-64M-C1
Text Generation
•
64.1M
•
Updated
Dec 30, 2025
•
24
aquilesfd/VLM-MoE-3x124M-Experiment
0.4B
•
Updated
Jun 18, 2025
•
2