armand0e commited on
Commit
da651b9
·
verified ·
1 Parent(s): 419fb48

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +154 -5
README.md CHANGED
@@ -1,21 +1,170 @@
1
  ---
2
- base_model: armand0e/Qwen3.5-9B-Agent
3
  tags:
4
- - text-generation-inference
5
  - transformers
 
 
 
 
 
6
  - unsloth
7
  - qwen3_5
8
  license: apache-2.0
9
  language:
10
  - en
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
11
  ---
12
 
13
- # Uploaded finetuned model
14
 
15
  - **Developed by:** armand0e
16
  - **License:** apache-2.0
17
- - **Finetuned from model :** armand0e/Qwen3.5-9B-Agent
18
 
19
  This qwen3_5 model was trained 2x faster with [Unsloth](https://github.com/unslothai/unsloth) and Huggingface's TRL library.
20
 
21
- [<img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png" width="200"/>](https://github.com/unslothai/unsloth)
 
1
  ---
2
+ base_model: Qwen/Qwen3.5-9B
3
  tags:
 
4
  - transformers
5
+ - text-generation-inference
6
+ - coding agent
7
+ - agent
8
+ - code
9
+ - tools
10
  - unsloth
11
  - qwen3_5
12
  license: apache-2.0
13
  language:
14
  - en
15
+ datasets:
16
+ - TeichAI/claude-4.5-opus-high-reasoning-250x
17
+ - armand0e/badlogicgames-pi-mono-opus-filtered
18
+ - armand0e/kimi-k2.6-claude-code-traces
19
+ - TeichAI/Claude-Opus-4.6-Reasoning-887x
20
+ - armand0e/minimax-m3-claude-code-traces
21
+ - armand0e/claude-opus-4.8-pi-traces
22
+ ---
23
+
24
+ # Qwen3.5 9B Coder
25
+
26
+ This is a experimental finetune on a mix of many traces from many different models. Reasoning was left untouched.
27
+
28
+ Total train time: ~4 hours
29
+
30
+ ## Training Script
31
+
32
+ <details>
33
+ <summary>Training Script</summary>
34
+
35
+ ```py
36
+ import os
37
+ from unsloth import FastModel
38
+ import torch
39
+ from trl import SFTConfig, SFTTrainer
40
+ from teich import mask_data, prepare_data
41
+
42
+ MAX_SEQ_LEN = 32768
43
+ MODEL_NAME = "Qwen/Qwen3.5-9B"
44
+ OUTPUT_DIR = "/content/drive/MyDrive/Colab/outputs-qwen-tool-sft"
45
+ HUB_REPO_ID = "armand0e/Qwen3.5-9B-Coder"
46
+ HF_TOKEN = os.environ.get("HF_TOKEN", "")
47
+ CHAT_TEMPLATE_PATH = "qwen3.5-chat-template.jinja"
48
+
49
+ model, tokenizer = FastModel.from_pretrained(
50
+ model_name=MODEL_NAME,
51
+ max_seq_length=MAX_SEQ_LEN,
52
+ load_in_4bit=False,
53
+ load_in_8bit=False,
54
+ full_finetuning=False,
55
+ token=HF_TOKEN,
56
+ )
57
+
58
+ if CHAT_TEMPLATE_PATH:
59
+ with open(CHAT_TEMPLATE_PATH, "r", encoding="utf-8") as f:
60
+ custom_chat_template = f.read()
61
+ tokenizer.chat_template = custom_chat_template
62
+ if hasattr(tokenizer, "tokenizer") and tokenizer.tokenizer is not None:
63
+ tokenizer.tokenizer.chat_template = custom_chat_template
64
+
65
+ model = FastModel.get_peft_model(
66
+ model,
67
+ finetune_vision_layers = False, # Turn off for just text!
68
+ finetune_language_layers = True, # Should leave on!
69
+ finetune_attention_modules = True, # Attention good for GRPO
70
+ finetune_mlp_modules = True, # Should leave on always!
71
+
72
+ r = 32, # Larger = higher accuracy, but might overfit
73
+ lora_alpha = 32, # Recommended alpha == r at least
74
+ lora_dropout = 0,
75
+ bias = "none",
76
+ random_state = 3407,
77
+ )
78
+
79
+ train_dataset = prepare_data(
80
+ {
81
+ "qwen3.7-max": {
82
+ "source": "armand0e/qwen3.7-max", # stupid typo i made and now this model wasn't trained on the qwen3.7-max traces :(
83
+ },
84
+ "chat": {
85
+ "source": "TeichAI/claude-4.5-opus-high-reasoning-250x",
86
+ },
87
+ "opus-pi-agent": {
88
+ "source": "armand0e/badlogicgames-pi-mono-opus-filtered",
89
+ },
90
+ "kimi-k2.6-claude-code": {
91
+ "source": "armand0e/kimi-k2.6-claude-code-traces",
92
+ },
93
+ "chat-2": {
94
+ "source": "TeichAI/Claude-Opus-4.6-Reasoning-887x"
95
+ },
96
+ "minimax-m3-claude-code": {
97
+ "source": "armand0e/minimax-m3-claude-code-traces"
98
+ },
99
+ "more-opus": {
100
+ "source": "armand0e/claude-opus-4.8-pi-traces"
101
+ }
102
+ },
103
+ tokenizer,
104
+ split="train",
105
+ hf_token=HF_TOKEN,
106
+ chat_template_kwargs={"enable_thinking": False, "preserve_thinking": True},
107
+ max_length=MAX_SEQ_LEN,
108
+ oversized_policy="trim_followups",
109
+ tokenize=True,
110
+ strict=True,
111
+ )
112
+
113
+ trainer = SFTTrainer(
114
+ model=model,
115
+ tokenizer=tokenizer,
116
+ train_dataset=train_dataset,
117
+ eval_dataset=None,
118
+ args=SFTConfig(
119
+ dataset_text_field="text",
120
+ dataset_num_proc=1,
121
+ max_length=MAX_SEQ_LEN,
122
+ packing=False,
123
+ per_device_train_batch_size=1,
124
+ gradient_accumulation_steps=8,
125
+ warmup_steps= 5,
126
+ num_train_epochs=1,
127
+ learning_rate=2e-4,
128
+ logging_steps=1,
129
+ save_strategy="epoch",
130
+ save_total_limit=3,
131
+ optim="adamw_8bit",
132
+ weight_decay=0.01,
133
+ #max_grad_norm=0.3,
134
+ lr_scheduler_type="linear",
135
+ output_dir=OUTPUT_DIR,
136
+ seed=3407,
137
+ report_to="none",
138
+ ),
139
+ )
140
+
141
+ trainer = mask_data(
142
+ trainer,
143
+ tokenizer=tokenizer,
144
+ train_on_reasoning=False,
145
+ train_on_final_answers=True,
146
+ train_on_tools=True,
147
+ )
148
+
149
+ print(trainer.train_dataset.preview())
150
+
151
+ trainer_stats = trainer.train(resume_from_checkpoint=False)
152
+
153
+ model.push_to_hub(f"{HUB_REPO_ID}-LoRA", token=HF_TOKEN)
154
+ tokenizer.push_to_hub(f"{HUB_REPO_ID}-LoRA", token=HF_TOKEN)
155
+
156
+ model.push_to_hub_merged(HUB_REPO_ID, tokenizer, save_method="merged_16bit", token=HF_TOKEN)
157
+ ```
158
+ </details>
159
+
160
  ---
161
 
162
+ The data for this model was easily formatted and masked with [Teich](https://github.com/TeichAI/teich)
163
 
164
  - **Developed by:** armand0e
165
  - **License:** apache-2.0
166
+ - **Finetuned from model :** Qwen/Qwen3.5-9B
167
 
168
  This qwen3_5 model was trained 2x faster with [Unsloth](https://github.com/unslothai/unsloth) and Huggingface's TRL library.
169
 
170
+ [<img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png" width="200"/>](https://github.com/unslothai/unsloth)