Vizuara AI Labs · instruction fine-tune

SLM‑125M  Instruct

Our 125M legal model, taught to follow tasks: summarize, extract, rewrite, classify, draft, and obey format constraints, on text you provide or from general domain knowledge.

125M
parameters
6,461
instructions
10
task types
6.85
val ppl
Validation metrics along the 125M lineage
Each perplexity is measured on that stage's own validation set, so read the trend as 'how well the model fits its own stage's data', not as one curve on one dataset. DPO and RLAIF optimize preferences rather than likelihood, so they log preference margin and reward instead of perplexity. Click a stage to open that model.
Base
ppl 8.36
pretrain val
QA SFT
ppl 6.06
QA val
Instruct
ppl 6.85
instruction val
DPO
margin 75.4%
preference val, no ppl
/
RLAIF
reward 9.9→11.7
RM reward, no ppl
RAFT on DPO
ppl 2.01
RAFT val
/
RAFT on RLAIF
ppl 2.04
RAFT val
instruction or question
optional: text to work on (attached as TEXT)
ready
The response will appear here.

What this is instruction SFT

An instruction-tuned stage of the 125M: the closed-book QA model was fine-tuned on ~6.5k domain-grounded synthetic instructions (summarize, extract, rewrite in plain English, classify, explain, draft, enumerate, format-constrained answers), every example compliance- and groundedness-judged. Lineage: base → QA SFT → instruction SFT.

Served scale-to-zero on Modal, so the first request may take ~20–60s while the model wakes.