Mati is a small language model fine-tuned to act as a math tutor for children (ages ~6–12). The core idea is simple but strict: Mati never solves the problem for the child — it teaches the method.
🧮 Try it live: the trilingual demo (català · español · English) runs free at huggingface.co/spaces/mediacloner/mati-tutor — CPU-only hosting, so a reply can take up to a minute. Want it on your own machine? See RUN_LOCAL.md.
Most chatbots happily hand over the answer to "What is 24 + 18?". That's terrible for learning. Mati is trained to follow three rules instead:
- Guide, don't solve — when asked a math question, Mati explains the procedure step by step ("add the units first, carry the one...") and invites the child to compute it themselves. The final answer is never revealed, even if the child insists ("just tell me!!").
- Verify answers — when the child proposes a result ("Is 24 + 18 = 42?"), Mati confirms or denies it. If it's wrong, Mati hints at where the mistake probably is — still without giving the answer.
- Everything comes back to math — off-topic questions get gently redirected into a fun math mini-challenge ("A T-Rex was 12 meters long, and you're about 1.5 — how many of you make one T-Rex?").
- Base model: Gemma-2-9B-it, fine-tuned with QLoRA (PEFT + TRL) on a single local RTX 3060 (12 GB). The tutor speaks Catalan, Spanish and English — one LoRA adapter per language, with the Spanish/English corpora generated as spec-driven ports of the Catalan one (
scripts/gen_dialogs.py). - Dataset: ~1,300 synthetic chat examples across six categories (method guidance, answer-verification correct/wrong, jailbreak resistance, off-topic redirection, multi-turn tutoring), with arithmetic ground truth generated programmatically so it's guaranteed correct.
- Calculator tool (v3+): a 9B model is an unreliable mental calculator, so Mati doesn't do the arithmetic itself. When the child proposes an answer, Mati emits a hidden
<calc>call, a sandboxed evaluator computes the exact result, and Mati narrates the verdict. Verification is checked against a real computation, never the model's head — and on a wrong answer the tool withholds the true value, so it can never slip out. - Evaluation: a hand-written 60-prompt test set plus scripted multi-turn chat probes, LLM-judged, measuring leak rate, verification accuracy, redirect rate, and resistance to insistence — always compared against the prompt-only baseline.
Final model (v8, with the calculator tool), LLM-judged on the 60-prompt test set vs the prompt-only baseline:
| What it must do | Baseline (prompt-only) | Mati v8 |
|---|---|---|
| Guide, don't solve — never leak the answer (Rule 1) | 0/20 leaks | 0/20 leaks |
| Verify answers correctly (Rule 2) | 1/20 — dodges everything | 20/20 |
| …and never reveal the answer when denying | — | 0/10 leaks |
| Resist "just tell me!!" (Rule D) | 10/10 | 10/10 |
| Redirect off-topic to math (Rule 3) | 9/10 | 9/10 |
| Catalan-language quality | 28/60 defective | ~1/60 |
The tool fires on 100% of verification turns and 0% of everything else, extracts the right expression 100% of the time, and behaves identically across multi-turn conversations.
The Spanish (es-v2) and English (en-v3) adapters, evaluated on ported 60-prompt test sets, match the Catalan results: 0/20 leaks, 20/20 verification, same tool behavior (see results/).
The interesting part wasn't the fine-tune — it was discovering what fine-tuning couldn't fix, and routing around it.
- The prompt-only baseline already refused to hand over answers and resisted insistence — but it dodged verification (1/20; it praised or deflected instead of judging true/false) and wrote defective Catalan (28/60).
- Supervised fine-tuning (v1–v2) turned the dodging into committed true/false verdicts (17/20) and cleaned up the Catalan — but verification then plateaued at 17/20 across two dataset iterations. Every residual miss was the model's own arithmetic failing (borrow-subtractions, a fraction). That's a 9B capacity limit; more data didn't move it.
- The calculator tool (v3–v6) took arithmetic off the model's plate. v3 lifted single-turn verification to 20/20 (all three v2 arithmetic failures fixed); v4 extended the tool to mid-conversation verification in multi-turn dialogs; v5 made the tool's feedback leak-proof (it withholds the true value on a wrong answer); v6 added a leak-safe direction hint so the tutor also tells the child which way their error went ("you're a bit short" / "a bit over") without revealing by how much. Verification is solved and the hints point the right way — single- and multi-turn, with zero answer leaks.
- The correct-Catalan prompt (v8) closed the last language gap. v6's Catalan was clean but not perfect (~3/60 residual — a mangled irregular verb, a dropped ela geminada), and those slips only surfaced in free-form generation like off-topic redirects. v8 folds an explicit correct-Catalan instruction into the system prompt and retrains so the fix is learned rather than nudged at inference; residual defects drop to ~1/60 with verification, resistance and leak-safety all unchanged. Same recipe as v6 otherwise (Gemma-2-9B + calculator tool); v7 was a separate Gemma-3 experiment that didn't beat v6.
The one item left is the exact size of the error — Mati can say you're too low but not "by ten" — and that's deliberate: stating the magnitude would let a child back out the answer. Naming the error size is out of reach without reopening the leak, so that's where the tool design stops.
Why fine-tune (and add a tool) instead of just a system prompt? Partly because tuned behavior is far harder for a persistent kid to override — and partly because that's the whole point: a hands-on, end-to-end exercise in shaping a small model's behavior, finding its limits, and engineering past them.
See DESIGN.md for the full story: dataset composition, the tool protocol, training config, per-version results (§6), failure analysis, risks, and roadmap. Per-run transcripts and metrics are in results/.
