This week on Reddit, a developer tried replacing Claude with a 287M-parameter model (less than 1/500th of GPT-4) to process 5 million German court rulings. The conclusion: it runs, but accuracy lags behind the teacher (Claude).
What this is
The essence of this post is a cost experiment in enterprise document processing.
The developer first had Claude Sonnet 'grade' 700 rulings — labeling out the names, institutions, legal clauses, amounts, and dates within them (this step is called 'data labeling'). They then used these labels to train a small model called GLiNER (a lightweight text-labeling AI, only a few hundred MB in size), teaching it to grade the remaining 4.99 million on its own.
The upside of large models is intelligence; the downside is that every processed document incurs API call fees (the cost of invoking an AI service each time), and 5 million of them adds up to astronomical figures. A small model, once trained, can be reused unlimited times at near-zero per-document cost — and this is the core question enterprise IT actually cares about: how to make AI affordable.
The developer's conclusion was honest: the small model can do the work, but it's always a bit behind the teacher. Nobody on Reddit could tell them exactly where it falls short.
How the industry sees it
The optimistic view is that this path is correct. Microsoft, Mistral, and HuggingFace have all been pushing 'small language models' (SLMs) for the past two years, arguing that enterprise AI deployment doesn't need to call GPT-4 every time. A 287M model can fit on an ordinary server, making it naturally attractive to industries with strict data compliance requirements (healthcare, legal, finance).
But this developer also raises a pragmatic counterargument: what he fears most is not 'the model isn't accurate enough' but 'I don't know where it's inaccurate.' Discovering a critical field was extracted wrong only after processing 5 million documents — in legal contexts, this risk is irreversible. So even with tempting economics, in high-stakes industries, that 'last 5% accuracy gap' may determine whether this approach can actually reach production.
There's another rarely discussed hidden risk: the training data itself was labeled by the large model, meaning the small model inherits all the biases and blind spots of the large model. Nobody in the Reddit comments gave this developer a clear answer either — which itself suggests this path hasn't yet been made to work.
What this means for regular people
For enterprise IT: if your company has large volumes of structured documents (contracts, invoices, customs declarations, medical records, court rulings), the 'large model labels data + small model runs production' path is worth a small-scale pilot. This isn't a technology problem — it's an engineering problem: whether you can absorb that last 5% of error.
For individual careers: people doing repetitive work like auditing, compliance, legal assistance, and document entry will be the first squeezed by these tools. It's not 'being replaced by AI' but 'being replaced by colleagues assisted by AI' — their workload and per-unit output will change significantly.
For consumer markets: consumers won't notice much in the short term. But the banks, insurance companies, and court systems you deal with may process your applications significantly faster and cheaper in three years — provided they're willing to absorb that last 5% of error.