Yeco Flash Mini-IT: The Italian Micro-LLM That Challenges Silicon Valley at Zero Cost
118M parameters, trained for under €15, ~200 tok/s on a consumer CPU. Yeco Flash Mini-IT is our proof of concept that sovereignty beats scale.
Why probabilistic LLMs need deterministic cognitive layers.
Marco Nasi
Founder & CEO
The current generation of Large Language Models (LLMs) has demonstrated remarkable capabilities in pattern matching, code generation, and natural language understanding. However, as we move from "chatbots" to "autonomous agents," a fundamental limitation has become increasingly apparent: probabilistic systems cannot guarantee deterministic behavior over long horizons.
LLMs operate on probability distributions. When an agent is tasked with a multi-step objective—say, "deploy this application to AWS"—it must maintain a coherent chain of thought across dozens of actions. A single low-probability token sample can derail the entire process, leading to what we call "Semantic Drift."
This is where the Context Window fails us. Simply stuffing more context into a model does not solve the problem of attention degradation. In fact, longer contexts often increase the noise-to-signal ratio, making hallucinations more likely, not less.
At YecoAI, we believe the solution lies not in bigger models, but in better architectures. The Anti-Loop Layer represents a shift towards "Cognitive Architectures"—systems where the LLM is treated as a reasoning engine, but not the executive controller.
A Cognitive Layer acts as an exogenous monitor. It sits outside the LLM, observing its inputs and outputs. It maintains a state that is independent of the model's context window.
Running a second GPT-4 to monitor your first GPT-4 is cost-prohibitive and slow. This is why our Anti-Loop Layer is built on Mini-LLMs—highly specialized, quantized models that can run on CPU with minimal latency.
By offloading the "executive function" to a lightweight, deterministic layer, we can achieve reliability that far exceeds what a raw LLM can provide, at a fraction of the cost.
As we continue to develop tools like Anti-Loop Layer and YecoAI Lab, our focus remains on this intersection of probabilistic power and deterministic control.
118M parameters, trained for under €15, ~200 tok/s on a consumer CPU. Yeco Flash Mini-IT is our proof of concept that sovereignty beats scale.
Ender-1 is our proprietary 3B model trained on EnderDevelopment data (free users, policy-compliant). It is in beta and rolling out on EnderDevelopment today.
278M-parameter neural model for PII detection and anonymization. Native Italian, 6 European languages, average F1 0.966.