YecoAI

Yeco Flash Mini-IT v1

118M-parameter Italian micro-LLM, trained for under €15, ~200 tok/s on CPU. Beats 5× larger models on Italian logic benchmarks.

Abstract

Yeco Flash Mini-IT is a proof-of-concept proprietary Italian LLM built and trained entirely in-house. While frontier labs burn billions on trillion-parameter models, this 118M-parameter architecture (118.1M measured) shows that hyper-specialized code and clean proprietary datasets matter more than brute force.

Efficiency record

  • 118M parameters (GPT-2 decoder-only + Context Gating, Pre-LN, weight tying)
  • Trained for under €15 total (pretraining ~$8.75 on H100 + SFT ~$4)
  • Runs locally on CPU: ~200 tokens/second on old consumer hardware
  • Custom Italian BPE 32k tokenizer, fertility 1.53 tok/word (−34% vs GPT-2)

Benchmarks (measured, ITA-Bench)

Scores are measured on official ITA-Bench datasets (SapienzaNLP), log-likelihood per choice, N=200/task, seed 42 — not estimated.

  • ARC-Challenge: 24.0% (Minerva-350M paper: 24.6%)
  • HellaSwag: 27.5% (Minerva: 32.6%)
  • BoolQ: 62.0% — beats Minerva-350M (60.7%)
  • TruthfulQA: 34.5% · GSM8K: 32.0% · Average: 36.0%

Honest reading: on multiple-choice benchmarks V2 is within statistical noise (±3.5 points/task) of the pre-SFT checkpoint. V2's value is behavioral: reliable question→answer→stop format, echo eliminated, structured chat answers.

License

Apache 2.0. Subject to rework: the model is under active iteration and numbers may change with future releases.

Related content