<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en-US"><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://www.albertsikkema.com/feed.xml" rel="self" type="application/atom+xml" /><link href="https://www.albertsikkema.com/" rel="alternate" type="text/html" hreflang="en-US" /><updated>2026-07-09T15:39:38+00:00</updated><id>https://www.albertsikkema.com/feed.xml</id><title type="html">Albert Sikkema - Building Production AI Systems</title><subtitle>Production-ready AI implementation, software engineering best practices, and enterprise AI systems development. Building scalable AI solutions with Claude, OpenAI, and engineering discipline for enterprise and government.</subtitle><author><name>Albert Sikkema</name></author><entry><title type="html">The Gap Is Not That Wide</title><link href="https://www.albertsikkema.com/ai/opinion/2026/06/28/the-gap-is-not-that-wide.html" rel="alternate" type="text/html" title="The Gap Is Not That Wide" /><published>2026-06-28T06:00:00+00:00</published><updated>2026-06-28T06:00:00+00:00</updated><id>https://www.albertsikkema.com/ai/opinion/2026/06/28/the-gap-is-not-that-wide</id><content type="html" xml:base="https://www.albertsikkema.com/ai/opinion/2026/06/28/the-gap-is-not-that-wide.html"><![CDATA[<figure>
  <img src="/assets/images/the-gap-is-not-that-wide.jpg" alt="Mind the gap warning painted in yellow on a London Underground platform edge" width="2592" height="1936" fetchpriority="high" style="width:100%;height:auto" />
  <figcaption>Photo by <a href="https://www.pexels.com/@pixabay">Pixabay</a> on <a href="https://www.pexels.com">Pexels</a></figcaption>
</figure>

<p>Every time I read that Europe will “never catch up” in AI, it annoys me. The distance is too large, the gap is permanent, we should just accept American dominance and move on. I keep hearing it from commentators, from tech media, from people who should know better.</p>

<p>It is nonsense.</p>

<p>We have a large population of highly educated, intelligent people. We have a culture that, for all its flaws, raises people to think about ethics, about climate, about responsibility toward fellow humans. Yes, we need to live up to those values more than we do. But we also massively underestimate ourselves. There is this persistent European reflex that everything is possible in America and that you have to move there to “make it” (whatever that means: getting extremely rich while ignoring the people sleeping in the street meters from your car, probably). If you think about it, that is not what most Europeans want. Too selfish, too wasteful, too careless about people and planet.</p>

<p>The premise that we are hopelessly behind is not supported by anything except the marketing of the companies that happen to be leading right now.</p>

<h2 id="i-have-switched-models-several-times">I Have Switched Models Several Times</h2>

<p>If we go back a long way (it feels like that, but it is only about 3,5 years) I started with GPT-3.5 for coding when it launched in late 2022. After that GPT-4. Then 5, then Opus. Every time I switched, the new model was better than the last, but never as good as marketed. Every time, the model I left behind had been called “the one that will never be surpassed” at launch. And every time it was said that developers would be out of a job in the next 6 months. None of this was true.</p>

<p>The <a href="https://artificialanalysis.ai/leaderboards/models">Artificial Analysis leaderboard</a> (updated regularly, and any snapshot is already outdated by the time you read this) tells a consistent story: the distance between American proprietary models and open-source alternatives is much smaller than the marketing suggests. Models that launched a year ago are routinely surpassed by competitors that barely existed at the time.</p>

<h2 id="why-does-the-hopeless-gap-story-persist">Why Does the “Hopeless Gap” Story Persist?</h2>

<p>So why does the “hopeless gap” story persist? Because it serves everyone who tells it.</p>

<p><a href="https://letsdatascience.com/news/openai-weighs-delaying-ipo-to-2027-3febe186">OpenAI</a> and <a href="https://www.cnbc.com/2026/06/01/anthropic-ipo-s1-prospectus.html">Anthropic</a> both filed S-1s (the registration to become publicly traded on the stock market) within a week of each other in early June 2026. OpenAI is targeting a valuation of up to a trillion dollars. When you are about to go public at those numbers, being mysterious about your capabilities and maintaining the narrative that nobody can touch you is not just marketing, it is a financial necessity. “We are so far ahead nobody can catch us” is the pitch, and every headline about export restrictions and “too dangerous” models reinforces it.</p>

<p>The <a href="https://fortune.com/2026/06/13/anthropic-disables-fable-mythos-export-controls-national-security-threat/">export controls</a> are a neat arrangement where everybody wins. Politicians get to look patriotic, protecting American interests and national security (and probably making some money along the way, given this administration’s track record with conflicts of interest). Companies get free headlines about how powerful their model must be if the government is scared of it, which is great for your stock price when you are about to go public. The cycle is predictable: a new model launches, it gets <a href="https://www.cnbc.com/2026/06/26/us-government-anthropic-claude-mythos5-ai.html">restricted</a> as “too dangerous,” and then restrictions quietly loosen just in time for the next model to take its turn. Media repeats the narrative because “permanent American dominance” is a better headline than “incremental progress across multiple geographies.”</p>

<p>And the anxiety it creates in Europe is real: we believe the marketing from those huge American companies and ‘slikken het als zoete koek’ (swallow it like candy, as we say in Dutch). The gap is there, but the perceived width of the gap is not real, and it is becoming a self-fulfilling prophecy.</p>

<h2 id="what-the-ai-timeline-actually-shows">What the AI Timeline Actually Shows</h2>

<p>If you zoom out even slightly, the pattern is obvious. This field is very, very young: about four years old in any practical sense. In those four years:</p>

<ul>
  <li>Google went from “their AI is embarrassing” to a serious contender in under two years</li>
  <li>Anthropic went from “ChatGPT knockoff” to having models <a href="https://fortune.com/2026/06/13/anthropic-disables-fable-mythos-export-controls-national-security-threat/">restricted by the US government</a> in three years</li>
  <li>DeepSeek went from unknown to <a href="https://api-docs.deepseek.com/news/news250120">shaking markets</a> in about a year</li>
  <li><a href="https://mistral.ai/">Mistral</a> went from a French startup to a model provider aiming to compete with the big three</li>
</ul>

<p>None of this is consistent with the idea of a permanent, uncatchable lead. It is consistent with a young, fast-moving field where leadership rotates and spending advantages erode quickly as the technology commoditises.</p>

<p>I <a href="/ai/opinion/cloud/2026/06/28/the-third-option-for-enterprise-ai.html">wrote earlier today</a> about European cloud providers running open-source models as a practical alternative to US hyperscalers. That post was about the solution. This one is about the fear that makes people think they do not need one: the idea that the race is already over.</p>

<p>It is not. The race barely started, and the starting positions change every year. If one to two years of lag were truly fatal, most of the models we use today would never have existed. And more exciting: what is possible in a year? How do we create and use these models in an ethical and compassionate way?</p>

<p><em>Working on AI strategy and wondering how the competitive field affects your choices? <a href="#" onclick="task1(); return false;">Get in touch</a> to compare notes.</em></p>

<h2 id="sources">Sources</h2>

<ul>
  <li><a href="https://artificialanalysis.ai/leaderboards/models">Artificial Analysis LLM Leaderboard</a> – Independent model comparison, updated regularly</li>
  <li><a href="https://fortune.com/2026/06/13/anthropic-disables-fable-mythos-export-controls-national-security-threat/">Fortune: Anthropic disables Fable/Mythos</a> – Export control whiplash</li>
  <li><a href="https://www.cnbc.com/2026/06/26/us-government-anthropic-claude-mythos5-ai.html">CNBC: Mythos 5 restrictions loosened</a> – Two weeks later</li>
  <li><a href="https://apnews.com/article/ai-data-centers-water-consumption-drought-arid-climate-f7f0727ca73f2b1e4ef2e51d90e45b0b">AP News: AI data centres and water consumption</a> – The resource cost of the spending race</li>
</ul>

<h2 id="related-posts">Related Posts</h2>

<ul>
  <li><a href="/ai/opinion/cloud/2026/06/28/the-third-option-for-enterprise-ai.html">The Third Option for Enterprise AI</a> – The practical alternative to US hyperscalers</li>
  <li><a href="/ai/opinion/research/2026/06/24/why-llms-will-not-have-your-next-big-idea.html">Why LLMs Will Not Have Your Next Big Idea</a> – More on separating hype from capability</li>
</ul>]]></content><author><name>Albert Sikkema</name></author><category term="ai" /><category term="opinion" /><summary type="html"><![CDATA[The 'can never catch up' AI narrative is contradicted by the industry's own timeline. A practitioner's perspective on why the anxiety is overblown.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://www.albertsikkema.com/assets/images/the-gap-is-not-that-wide-blog.png" /><media:content medium="image" url="https://www.albertsikkema.com/assets/images/the-gap-is-not-that-wide-blog.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">The Third Option for Enterprise AI</title><link href="https://www.albertsikkema.com/ai/opinion/cloud/2026/06/28/the-third-option-for-enterprise-ai.html" rel="alternate" type="text/html" title="The Third Option for Enterprise AI" /><published>2026-06-28T00:00:00+00:00</published><updated>2026-06-28T00:00:00+00:00</updated><id>https://www.albertsikkema.com/ai/opinion/cloud/2026/06/28/the-third-option-for-enterprise-ai</id><content type="html" xml:base="https://www.albertsikkema.com/ai/opinion/cloud/2026/06/28/the-third-option-for-enterprise-ai.html"><![CDATA[<figure>
  <img src="/assets/images/third-option-enterprise-ai.jpg" alt="European countryside with dramatic clouds over rolling hills" width="1920" height="1280" fetchpriority="high" style="width:100%;height:auto" />
  <figcaption>Photo by <a href="https://unsplash.com/@adlan">Adlan</a> on <a href="https://unsplash.com">Unsplash</a></figcaption>
</figure>

<p>Almost all AI services I use with the different companies and governments I worked with are American based, mostly Azure and Anthropic. As we are all aware this is not a situation that is acceptable: especially not with ‘the orange child in the white house’ spewing nonsense every day and keeping us on our toes, American companies are no longer a trustworthy partner. They obviously chose sides during the orange child’s election, marking a departure from the heralded status American tech companies had over the last decades.</p>

<p>So all kinds of solutions are proposed to solve this. And basically there are two main trains of thought (or perhaps three). Option one: keep using the big American cloud providers (OpenAI, Anthropic, Google) and accept that your data crosses the Atlantic. As you can derive from the previous paragraph this is technically an option, but in practice definitely not: trust has been broken and will take many years to rebuild (if at all, given the current direction of American foreign politics). The alternative that I get to see pitched a lot: go fully local, run your own hardware, keep the data in the building. (Could be an option, I was in that camp for a long time, but what I do not understand about this proposition is that those who propagate it, do not mention that this would only be logical if we would move back from the cloud to full on-prem hosting of all data. For various reasons I do not see this happening anytime soon)</p>

<p>So that leaves a third option that makes more practical sense to me and to most European companies and is a big business opportunity within Europe: hosting opensource LLMs on European servers by European companies, that make sure privacy and data governance are guaranteed.</p>

<h2 id="option-1-stay-on-us-hyperscalers">Option 1: Stay on US Hyperscalers</h2>

<p>The current situation is using US companies for LLM: it works, it is what we are used to. However this is not sustainable, both in costs and in privacy and risking being cutoff of the service on a whim of the orange child.</p>

<p>Given that it is hard to filter PII or medical or legal data from a prompt, it is a minefield of GDPR and privacy issues that you need to navigate. And that is assuming you are able to make sure your employees only use your approved methods (do not forget your employees are infinitely creative to work around them, so you better make the approved methods really easy to use and perform really well, otherwise you could as well not have bothered at all)</p>

<h2 id="option-2-go-fully-on-prem">Option 2: Go Fully On-Prem</h2>

<p>The counterargument is to bring everything in-house. Buy GPUs, rack them in your server room, run open-source models, keep all data on your own hardware. Full control, full sovereignty, zero legal risk.</p>

<p>The edge AI market is real and growing, and there are genuine use cases for local inference: air-gapped environments, defence, certain healthcare scenarios.</p>

<p>But here is where it gets illogical for most companies. These same organisations run their email on Microsoft 365, their CRM on Salesforce, their databases on Azure or AWS, their file storage on SharePoint. They trust cloud providers with all of that. And then the argument is that specifically the LLM part, the thing that processes text and returns text, needs to be on physical hardware in the basement?</p>

<p>If your data governance concern is serious enough to run AI on-prem, it is serious enough to move everything back on-prem. And that is probably not going to happen.</p>

<p>There is also the cost problem. Running inference locally requires expensive GPUs, it is really complicated, so you need people who know how to operate them, and continuous investment to keep up with model improvements. For companies that are not in the AI infrastructure business, this is a distraction. (even though personally I believe every company should be in full control of its own data and software and I know this belief is shared by many IT and data professionals, this is unfortunately not a shared vision across most boards)</p>

<h2 id="option-3-european-providers-open-source-models">Option 3: European Providers, Open-Source Models</h2>

<p>The option I find most interesting is getting discussed more and more, but not yet seriously enough: European cloud providers running open-source models under European jurisdiction.</p>

<p>The landscape here has changed dramatically in the last two years. <a href="https://mistral.ai/">Mistral AI</a> now offers models under Apache 2.0 with inference running entirely in EU data centres. <a href="https://www.scaleway.com/">Scaleway</a> offers managed GPU instances with model-as-a-service APIs under French legal jurisdiction. <a href="https://www.t-systems.com/de/en/insights/newsroom/news/ai-sovereignty-for-germany-and-europe-1124980">Deutsche Telekom invested over a billion euros</a> in an Industrial AI Cloud with 10,000 GPUs in Munich. <a href="https://www.ionos.com/">IONOS</a> offers dedicated GPU servers. Even smaller players like <a href="https://regolo.ai/">Regolo.ai</a> provide GDPR-compliant inference with zero data retention policies.</p>

<p>And the models are good enough. Open source models compete with proprietary models on most practical tasks. For the “boring” enterprise use cases that make up 90% of actual AI adoption (document processing, summarisation, classification, translation, code assistance), open-source models hosted in Europe would do the job.</p>

<p>And costs of inference are dropping rapidly: <a href="https://epoch.ai/data-insights/llm-inference-price-trends">Epoch AI’s research</a> shows inference costs dropped 1,000x between 2021 and 2026 for GPT-3 level performance. The median decline is 50x per year, accelerating to 200x per year post-2024. GPU compute cost follows <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6198738">Wright’s Law</a> with an 89% learning rate per doubling of cumulative production. That is faster than solar panels ever dropped (around 20% learning rate).</p>

<h2 id="why-this-makes-sense">Why This Makes Sense</h2>

<p>The logic is simple. If you already trust cloud providers to run your email, your databases, your file storage, and your business applications, then the question is not “cloud versus on-prem.” The question is “which cloud provider, under which legal jurisdiction.”</p>

<p>A European provider running opensource models under EU law gives you:</p>

<ul>
  <li>No Chapter V transfer headaches</li>
  <li>No CLOUD Act exposure</li>
  <li>Full <a href="https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai">EU AI Act</a> compliance (high-risk obligations enforceable from August 2026, with extended deadlines into 2027-2028)</li>
  <li>The same operational model your IT team already knows (APIs, managed services, SLAs)</li>
</ul>

<p>The market agrees: <a href="https://www.intelligentcio.com/eu/2026/02/04/europe-accelerates-shift-to-region-specific-ai-platforms-amid-sovereignty-push/">Gartner reports</a> that 61% of Western European CIOs plan to increase reliance on local cloud providers, and European sovereign cloud spending is projected to grow 83% year-over-year to $12.6 billion in 2026. The <a href="https://www.cnbc.com/2026/04/24/cohere-aleph-alpha-germany-ai-europe-expansion.html">Cohere acquisition of Aleph Alpha</a>, backed by $600 million from Schwarz Group, signals that serious money is flowing into this space.</p>

<h2 id="when-this-does-not-work">When This Does Not Work</h2>

<p>Frontier reasoning tasks. If you need the absolute best model for complex multi-step reasoning, agentic coding, or advanced multimodal work, US providers still lead. The gap is closing, but it is definitely still there. As a rule of thumb: the best opensource models are about a year behind the best closed source models. Last January I noticed Opus was at such a level that I could actually start using it as I hoped, so it is reasonable to assume that coming January I will be able to do the same with an opensource model: exciting! An interesting complication is the <a href="https://www.washingtonpost.com/business/2026/06/26/trump-ai-openai-gpt56-sol-cybersecurity-mythos/cdd6f804-7181-11f1-8730-e7fd0e2a6404_story.html">recent US government blocking</a> of Fable 5 and GPT-5.6. If the American government keeps pulling its own models from the market, the “first mover” advantage disappears, and investment in open-source alternatives accelerates. But that is a topic for another post. Given the year advantage, we will be having that same functionality in an opensource model in June 2027 anyway.</p>

<p>For most enterprise workloads though, “the best model” is not the relevant criterion. “Good enough, legally clean, and operationally simple” is. And that is where European providers are becoming competitive.</p>

<p>The future I see is not a binary choice between American clouds and on-prem hardware. It is trustworthy European partners that run the LLM part, just as well as the other cloud services a company needs. And as a benefit to keep going that path: the European rules are among the toughest in the world, if a company can meet them here, it is quite possible to offer these services in other parts of the world, like South America and Africa.</p>

<p><em>Thinking about where to run your enterprise AI? <a href="#" onclick="task1(); return false;">Get in touch</a> to talk through the options.</em></p>

<h2 id="sources">Sources</h2>

<ul>
  <li><a href="https://epoch.ai/data-insights/llm-inference-price-trends">Epoch AI: LLM inference price trends</a> – 1,000x cost reduction data</li>
  <li><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6198738">Wright’s Law extended to GPU compute</a> – 89% learning rate vs solar’s 20%</li>
  <li><a href="https://www.dataprivacyframework.gov/Program-Overview">EU-US Data Privacy Framework</a> – Current legal basis for transatlantic transfers</li>
  <li><a href="https://www.aigovhub.io/guides/gdpr-eu-us-data-transfers-post-schrems-ii-guide-2026">GDPR and EU-US data transfers guide 2026</a> – Post-Schrems II compliance landscape</li>
  <li><a href="https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai">EU AI Act</a> – Full enforcement August 2026</li>
  <li><a href="https://www.intelligentcio.com/eu/2026/02/04/europe-accelerates-shift-to-region-specific-ai-platforms-amid-sovereignty-push/">Gartner: Europe accelerates sovereign AI shift</a> – 61% of CIOs increasing local provider use</li>
  <li><a href="https://www.precedenceresearch.com/edge-ai-market">Edge AI market projections</a> – $25B (2025) to $143B (2034)</li>
  <li><a href="https://www.cnbc.com/2026/04/24/cohere-aleph-alpha-germany-ai-europe-expansion.html">Cohere acquires Aleph Alpha</a> – $20B combined valuation, sovereign AI focus</li>
  <li><a href="https://www.t-systems.com/de/en/insights/newsroom/news/ai-sovereignty-for-germany-and-europe-1124980">Deutsche Telekom Industrial AI Cloud</a> – EUR1B+, 10,000 GPUs</li>
  <li><a href="https://www.washingtonpost.com/business/2026/06/26/trump-ai-openai-gpt56-sol-cybersecurity-mythos/cdd6f804-7181-11f1-8730-e7fd0e2a6404_story.html">OpenAI and Anthropic limit new AI models to Trump-approved customers</a> – Washington Post on US export controls</li>
  <li><a href="https://tweakers.net/nieuws/249082/vs-blokkeert-anthropic-claude-fable-5-en-mythos-5-vanwege-zorgen-jailbreak.html">VS blokkeert Anthropic Claude Fable 5 en Mythos 5</a> – Tweakers (Dutch) on the same topic</li>
</ul>]]></content><author><name>Albert Sikkema</name></author><category term="ai" /><category term="opinion" /><category term="cloud" /><summary type="html"><![CDATA[Enterprise AI does not have to be US hyperscalers or on-prem. European providers running open-source models offer a third path that makes more practical sense.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://www.albertsikkema.com/assets/images/the-third-option-for-enterprise-ai-blog.png" /><media:content medium="image" url="https://www.albertsikkema.com/assets/images/the-third-option-for-enterprise-ai-blog.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Why LLMs Will Not Have Your Next Big Idea</title><link href="https://www.albertsikkema.com/ai/opinion/research/2026/06/24/why-llms-will-not-have-your-next-big-idea.html" rel="alternate" type="text/html" title="Why LLMs Will Not Have Your Next Big Idea" /><published>2026-06-24T00:00:00+00:00</published><updated>2026-06-24T00:00:00+00:00</updated><id>https://www.albertsikkema.com/ai/opinion/research/2026/06/24/why-llms-will-not-have-your-next-big-idea</id><content type="html" xml:base="https://www.albertsikkema.com/ai/opinion/research/2026/06/24/why-llms-will-not-have-your-next-big-idea.html"><![CDATA[<figure>
  <img src="/assets/images/llm-limitations-boundary.jpg" alt="Moody coastline with cliff edge disappearing into mist, boundary between land and sea" width="1920" height="1280" fetchpriority="high" style="width:100%;height:auto" />
  <figcaption>Photo by <a href="https://unsplash.com/@phillbrown">Phill Brown</a> on <a href="https://unsplash.com">Unsplash</a></figcaption>
</figure>

<p>I have been using LLMs daily for years now, mostly to code, but also for research, writing, and analysis. They are genuinely useful, sometimes spectacularly so. I have built tools, workflows, and entire coding pipelines around them. So this is not a “LLMs are bad” post. This is about something I have been thinking about for a while, and an enthusiastic youtube video about the <a href="https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f">LLM Wiki pattern</a> got me thinking again about the drawbacks (besides the obvious benefits) and why I think that way of gathering and using data is not good enough for me.</p>

<p>In this post I will take you through two ideas: the first one is ‘LLMs can not create truly new ideas’ and the second is that ‘the simplification of information in an early stage leads to suboptimal results’. And the end conclusion is that, although we can do magnificent things with LLMs, this technique is not able to solve any real problems and will always result in mediocrity.</p>

<h2 id="the-argument-from-mechanism">The Argument From Mechanism</h2>

<p>Next-token prediction works by sampling from a distribution fitted over the training corpus. There is no reasoning faculty that could exceed what the text encodes. The model stores a very high-dimensional interpolation function over token sequences. It produces outputs for inputs not exactly in the training set (that is what makes it useful), but it interpolates within the <a href="https://arxiv.org/abs/2110.09485">convex hull</a> of expressed human thought. It does not extrapolate beyond the manifold it was fitted to, and the model does not learn.</p>

<p>What I here define as a good idea is a genuine breakthrough that is good <em>because</em> it violates the regularities the corpus taught. That is structurally outside the region the model can reach. To the model, the revolutionary idea and the incoherent error look the same: both are low-probability under the learned distribution. There is no mechanism to favour “low-probability because true-and-new” over “low-probability because wrong.”</p>

<p>The model’s only notion of “good” is “high-probability under the learned distribution,” which means “resembles what was already thought.” That is anti-correlated with genuine novelty almost by definition. The quality filter is filtering for conventionality, which explains why every LLM output feels competent but unsurprising. Ask any LLM to build you a web app and you will get React, Tailwind, and shadcn/ui. And only if you are an expert (or have a good general knowledge) you will spot the things it misses or simply fails on (your knowledge on that domain is then above medium).</p>

<h2 id="the-argument-from-scale">The Argument From Scale</h2>

<p>Now you could argue that there are a lot of good solutions in a LLM (including the novel ones we need to solve important issues like climate, energy and water management). And that they are simply buried in the vast knowledge of the LLM and need to be combined over different domains and then extracted. But if recombinational novelty could produce breakthroughs, the current reality would have surfaced them. We have the largest idea-discrimination experiment ever running right now, with millions of LLM instances running daily, with humans reading and judging. If LLMs would live up to their promise (the big breakthrough for humanity speech we have been hearing for a few years now): why do we not see those novel ideas showing up?</p>

<p>The result of those LLMs has a different character: a flood of interpolative novelty (restatements, syntheses, transfers of known methods to adjacent problems) and nothing on extrapolative novelty (fundamental breaks, genuinely new ideas). That bimodal signature is exactly what the mechanism predicts. Abundant where the manifold is dense (between known points), absent where it is not (beyond them).</p>

<p>If the good recombinations were reachable but merely unrecognised, millions of human discriminators sifting the output over years would have found them. They have not. The breakthroughs are not hidden in the output. They are simply not there.</p>

<p>Recent research backs this up. <a href="https://arxiv.org/pdf/2504.12320">Chakrabarty et al. (2025)</a> found that LLM-generated ideas lead to homogenous outcomes across domains, and <a href="https://arxiv.org/pdf/2504.15266">Tian et al. (2025)</a> showed that aligned LLMs remain trapped in a “safe attractor basin” created during RLHF, limiting exploratory novelty even with strong creativity prompts. There is a <a href="https://arxiv.org/pdf/2412.14141">counterpoint from Jiang et al. (2024)</a> showing LLMs can combine existing concepts in ways rated more novel than ideas from NLP researchers, but that is combinatorial novelty (recombining known things), not extrapolative novelty (ideas that break the frame).</p>

<h2 id="the-pipeline-problem">The Pipeline Problem</h2>

<p>This connects to something I was reminded of while looking at Karpathy’s <a href="https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f">LLM Wiki pattern</a> and <a href="https://gist.github.com/rohitg00/2067ab416f7bbe447c1977edaaa681e2">rohitg00’s v2 extension</a>. A wiki is basically a collection of datapoints (documents, website, databases etc) that can be connected based on theme or subject. This leads to a great overview of data on a subject with beautiful graphs that can help you understand interconnectedness of the data points. However most wikis fail here: they compile raw sources into a structured, interlinked knowledge graph at ingestion time, resulting in a more or less fixed but definitely 1 to low-dimensional description of connections. Before any question has been asked, before anything is known about what can safely be dropped, irreversible choices are made at the point of minimum information about what those choices should preserve. And the problem here is that an LLM makes those choices, easy but at the cost of massive information loss.</p>

<p>This is not just a gut feeling. Information theory has a name for it: the <a href="https://en.wikipedia.org/wiki/Data_processing_inequality">Data Processing Inequality</a>. It says that processing data can only destroy information, never create it. If X goes through transformation T to produce Y, then Y cannot contain more information about the original source than X did. Every step in a pipeline can only lose or preserve, never gain. It is the same discipline as keeping the RAW file instead of the JPEG, or lazy evaluation in programming: defer the irreversible reduction until the moment you know what you actually need.</p>

<p>The wiki does the opposite. It flattens at ingestion, the earliest possible moment, when it has the least information about what any future question will require. When the target artifact is a node-edge graph, only graph-shaped information survives extraction. Argument structure, evidential weight, mechanism, scope conditions, the reasoning connecting premise to conclusion: gone. Not because the model failed, but because it succeeded at building exactly the thing it was aimed at, and that thing has no room for any of it. A wiki link (even a typed one like “refutes”) collapses a multi-dimensional relationship into a single undifferentiated connection. The real relationship between two claims has independent components: logical relation, evidential strength, domain, temporal validity. The graph has one slot.</p>

<p>For planning a holiday this does not matter, losing some nuance about hotel reviews may not be so important (and you can deliberately choose that this is not important). But for anything where the quality of reasoning matters (medical research, legal analysis, policy decisions, engineering trade-offs), early information loss is crippling. For example three pro-articles enter a pipeline; the compiled page reflects the pro-standpoint and see it as the base stance. Then a fourth article is added with a very strong point contra, that should weigh much heavier than the three pro documents with light evidence. The model has no faculty for evaluating that the fourth article’s methodology is stronger. Evidential weight and citation count are different quantities, and the system only has access to the latter. Worse, the three pro-articles arrived first, so they created the concept pages, set the framing, and named things. The dissenting source arrives into a wiki already shaped to express the opposite position. It gets architecturally demoted to a status of a light contra, basically a footnote, where it should dismiss the entire pro standpoint.</p>

<p>And then it starts adding up: in a serial pipeline (entity extraction, graph building, contradiction resolution, consolidation), each step consumes the previous step’s output as ground truth. Chain a 95%-reliable step ten times and you end up with a reliability of 60% (0.95^10 –&gt; 0.60). And on top of that these errors are correlated toward plausibility and majority framing, so the real number is worse. This is well-documented in cascading ML systems: <a href="https://papers.nips.cc/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html">Sculley et al. (2015)</a> at Google showed that ML pipelines accumulate “technical debt” where each model’s errors become the next model’s training signal, creating feedback loops that are extremely hard to detect or debug. LLMs are slightly different but it is reasonable to state that they are not better at this.</p>

<p>Making it extra complicated if you are not a domain expert is that perceived reliability and actual reliability diverge in opposite directions as the pipeline grows. Each added step makes the surface richer, more complete-looking, more authoritative (more scores, typed edges, structure) while simultaneously multiplying against the reliability floor. The more elaborate the system, the more trustworthy it looks and the less trustworthy it (most likely) is.</p>

<p>The correct approach to choosing where to cut information density is as late and as informed as possible. The cut should be applied at the query, shaped by the query, discarding only what that specific question does not need.</p>

<p>And this connects directly to the novelty problem. Every mechanism in these systems (confidence scoring, contradiction resolution, consolidation, forgetting curves) is tuned to suppress exactly the signature of genuine novelty: low support, high conflict, poor fit with existing structure. A good new idea is an anti-consensus object. The pipeline converts whatever novelty does appear into “noise to be cleaned up,” because it has no faculty to see it as anything else.</p>

<h2 id="semi-intelligence">Semi-Intelligence</h2>

<p>So we can define what we are actually dealing with: full competence within the convex hull of expressed human thought, zero reach outside it.</p>

<p>This has quite some implication: most practical work lives inside that hull: writing code, synthesising research, drafting documents, debugging, refactoring, transferring methods from one domain to an adjacent one. LLMs are quite good at all of this. I use them for it every day and they make me measurably more productive.</p>

<p>But “inside the hull” has a boundary, and everything interesting about the future of ideas happens at or beyond that boundary.</p>

<h2 id="the-great-equaliser">The Great Equaliser</h2>

<p>There is a way to frame this that makes the stakes clearer: the LLM is a nivellator. It pulls everyone toward the median of existing thought.</p>

<p>If you were below that median, you get lifted up. That is genuine value. I <a href="/ai/development/opinion/2026/06/15/why-diy-software-is-great-until-it-is-not.html">wrote about this before</a>: the democratisation of capability is real and good. Someone who could not write code can now build a working application. Someone who could not write a decent email can now produce clear, professional communication.</p>

<p>But if you were above that median, or thinking orthogonally to it, the LLM pulls you <em>down</em> toward consensus. Your unusual angle gets smoothed into the conventional take. Your weird framing gets normalised. Your instinct that the consensus is wrong gets buried under a fluent, well-structured argument for why the consensus is right. The tool does not argue with you. It just quietly steers everything toward what the training data says is most likely.</p>

<p>The net effect is convergence toward the mean. The better the model gets, the stronger the pull. More people producing more competent work, all of it sounding increasingly alike. Ask five people to use an LLM to write a strategy document and you get five documents that could have been written by the same person.</p>

<p>If you are building systems on top of LLMs (and I am), be honest about what they can and cannot do. They are navigation tools for existing knowledge, not idea generators. Do not expect an LLM pipeline to surface the insight that changes your approach. It will give you a very competent synthesis of what is already known, and that is genuinely valuable, but it is a different thing.</p>

<p>The people who matter most here are the ones the LLM is worst at emulating: the imaginary thinkers, the artists, the scientists, the outliers, the stubborn contrarians who look at the consensus and say “but what if that is wrong.” People whose judgment is not a statistical shadow of what has already been thought, but a faculty that can value the unfamiliar as such. They were always important, and in a world where the median is increasingly automated, they are irreplaceable.</p>

<p>So hurray to the unconventional thinkers: the world needs you, so keep thinking those original thoughts.</p>

<p><em>Thinking about what LLMs can and cannot do for your team? <a href="#" onclick="task1(); return false;">Get in touch</a> to compare notes.</em></p>

<h2 id="sources">Sources</h2>

<h3 id="on-llm-creativity-and-limitations">On LLM creativity and limitations</h3>

<ul>
  <li><a href="https://arxiv.org/pdf/2504.12320">Has the Creativity of Large-Language Models Peaked?</a> (Chakrabarty et al., 2025) – LLM-generated ideas lead to homogenous outcomes across domains</li>
  <li><a href="https://arxiv.org/pdf/2504.15266">Roll the Dice &amp; Look Before You Leap</a> (Tian et al., 2025) – Aligned LLMs stay trapped in RLHF’s “safe attractor basin”</li>
  <li><a href="https://arxiv.org/abs/2505.23323">Neither Stochastic Parroting nor AGI</a> (May 2025) – LLMs extrapolate from training priors but only within domain boundaries</li>
  <li><a href="https://arxiv.org/pdf/2412.14141">LLMs Can Realize Combinatorial Creativity</a> (Jiang et al., 2024) – Counterpoint: combinatorial novelty is real, but distinct from extrapolative novelty</li>
  <li><a href="https://arxiv.org/abs/2110.09485">Learning in High Dimension Always Amounts to Extrapolation</a> (Balestriero et al., 2021) – The convex hull problem in high-dimensional learning</li>
</ul>

<h3 id="on-information-loss-and-llm-pipelines">On information loss and LLM pipelines</h3>

<ul>
  <li><a href="https://en.wikipedia.org/wiki/Data_processing_inequality">Data Processing Inequality</a> – Information theory: processing can only destroy information, never create it</li>
  <li><a href="https://papers.nips.cc/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html">Hidden Technical Debt in Machine Learning Systems</a> (Sculley et al., 2015) – Google’s analysis of cascading errors in ML pipelines</li>
  <li><a href="https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f">LLM Wiki pattern</a> (Karpathy, 2026) – Compiled knowledge as alternative to RAG</li>
  <li><a href="https://gist.github.com/rohitg00/2067ab416f7bbe447c1977edaaa681e2">LLM Wiki v2</a> (rohitg00, 2026) – Extension with confidence scoring, temporal decay, and knowledge governance</li>
</ul>

<h3 id="related-posts">Related posts</h3>

<ul>
  <li><a href="/ai/development/tools/2026/04/16/when-llms-actually-deliver.html">When LLMs Actually Deliver</a> – Where LLMs shine: the interior of the hull</li>
  <li><a href="/ai/development/opinion/2026/06/15/why-diy-software-is-great-until-it-is-not.html">Why DIY Software Is Great Until It Is Not</a> – Similar theme: tools change what is possible, but not what requires expertise</li>
  <li><a href="/ai/development/opinion/2026/02/13/let-the-ai-pick-react.html">Let the AI Pick: React</a> – The conventionality filter in action: every LLM picks the same stack</li>
</ul>]]></content><author><name>Albert Sikkema</name></author><category term="ai" /><category term="opinion" /><category term="research" /><summary type="html"><![CDATA[After years of daily LLM use, I think we can draw a line: extraordinary tools for working within known territory, structurally unable to cross beyond it.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://www.albertsikkema.com/assets/images/why-llms-will-not-have-your-next-big-idea-blog.png" /><media:content medium="image" url="https://www.albertsikkema.com/assets/images/why-llms-will-not-have-your-next-big-idea-blog.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Why DIY Software Is Great Until It Is Not</title><link href="https://www.albertsikkema.com/ai/development/opinion/2026/06/15/why-diy-software-is-great-until-it-is-not.html" rel="alternate" type="text/html" title="Why DIY Software Is Great Until It Is Not" /><published>2026-06-15T00:00:00+00:00</published><updated>2026-06-15T00:00:00+00:00</updated><id>https://www.albertsikkema.com/ai/development/opinion/2026/06/15/why-diy-software-is-great-until-it-is-not</id><content type="html" xml:base="https://www.albertsikkema.com/ai/development/opinion/2026/06/15/why-diy-software-is-great-until-it-is-not.html"><![CDATA[<figure>
  <img src="/assets/images/diy-software-craftsmanship.jpg" alt="Old carpentry workshop with hand tools hanging on a wooden wall" width="1920" height="1280" fetchpriority="high" style="width:100%;height:auto" />
  <figcaption>Photo by <a href="https://unsplash.com/@ricky_kharawala">Ricky Kharawala</a> on <a href="https://unsplash.com">Unsplash</a></figcaption>
</figure>

<p>Most people I know put wallpaper up in their own house, do small to medium jobs in and around the house themselves and only hire out to a crafts(wo)man when the job is too big or takes too much time. But doing stuff yourself is more or less normalised: it saves money and is fun to do. The same is now true for software engineering: with AI and vibe coding, anyone can build something that works.</p>

<p>That is a genuinely great thing. The DIY movement is a democratising development that put real capability into the hands of ordinary people, created an entire industry around it, and gave millions of people the satisfaction of building something with their own hands.</p>

<p>I also had the pleasure of seeing some really experienced carpenters work up close (30+ years). An interesting observation I made is that they do not seem to work very hard at all. However when you come back an hour later, real progress is made. No movement is wasted, every step is known and the whole workflow is optimised to minimise waste of resources and time. They know in advance how much time a certain job will take, what they need to do it and thus what it will cost.</p>

<p>I have also friends that bought old, badly maintained houses and turned them into small palaces in a few years. There were some hurdles to take and a lot of lessons learned, but the result is beautiful.</p>

<p>So what is the difference? One is a professional, the other is not. But what does that mean? Professional in my book was always someone who got paid to do a job, an amateur could do the same but was not paid for it. However there is a deeper distinction to be made: the diy-er builds for themselves, while the professional builds for others (and gets paid for it)</p>

<p>When evaluating the advent of AI and the impact it has on software engineering you can probably already see the parallel I want to make. AI is a great tool that helps the diy-er to build something really nice, perfectly adapted to what they want. However building for others (professional software engineering) has other demands.</p>

<h2 id="the-software-version-of-this-story">The Software Version of This Story</h2>

<p>AI-assisted coding is doing for software what the bouwmarkt (home improvement store, for non-Dutch readers) did for building and home improvement. Not just the tools, but the tools and the advice and the materials, all in one place. An IDE or a linter is a power tool: it makes you faster at what you already know. AI is the whole store: you can walk in knowing nothing and walk out with a plan, the supplies, and enough confidence to get started. And I think that is great.</p>

<p>I <a href="/ai/development/2026/02/05/vibe-coding-quality-democratisation.html">wrote about this before</a>: the democratisation of software creation is genuinely exciting. The vibe-coded app that tracks your running schedule or manages your recipe collection or automates your invoicing? Build it. Have fun. Who cares if the code is messy. It works for you, and that is all that matters.</p>

<p>But then the app works so well that a friend wants it. Then a few more people want it. Someone suggests charging for it. And suddenly you are no longer building for yourself.</p>

<p>This is where it gets hard, and it is not about whether you can technically build the software. It is about everything else. When you build for others, you are working with someone else’s idea of what the product should do, not yours. You are responsible for their data, their money, their trust. You have to think about edge cases you would never hit yourself, because your users will hit all of them. You have to maintain it when you would rather be building something new. You carry financial risk if it breaks. You might need to comply with regulations you have never heard of.</p>

<p>More and more I get requests from diy-ers (non-developers, most of them in a company context) that ask me to evaluate their vibe-coded application. Most of the time the functionality is great (for their purposes), but there are a lot of open ends that need to be addressed.</p>

<p>Just as a professional carpenter would do when assessing a job, you need the expertise and the know-how (which goes far beyond the code itself, just as the carpenter knows far more than how to use a handsaw to saw wood) to analyse the situation and decide what is the best course of action: build on top of the application or tear it all down and start with a proper foundation.</p>

<h2 id="it-is-not-about-skill">It Is Not About Skill</h2>

<p>I have made the transition from diy to professional twice. Before software, I worked as a sound engineer. That also started as a hobby: years of messing around with equipment, doing small gigs, learning by doing. The moment I went professional, the job changed. It was no longer about what sounded good to me, it was about what the client needed, on deadline, within budget, with backup plans for when things went wrong. The technical skills were the same, but everything around them was different.</p>

<p>Then I did it again with software. I had been programming for over twenty-five years for myself before making the switch to professional software engineering. The code I wrote as a hobbyist was fine for what it was. But writing code that other people depend on, that needs to run reliably, that handles their data responsibly, that complies with regulations I had never thought about was another game.</p>

<p>So this is not diy being bad and professionals good. Plenty of professional software is terrible and plenty of hobbyist software is brilliant. The distinction is not about ability.</p>

<p>The distinction is about context. When you build for yourself, you optimize for one person’s needs, and you can cut every corner you want because you are the only one who bears the consequences. When you build for others a lot of this changes:</p>

<ul>
  <li>It is not your use case anymore, so you have to understand problems you do not personally have</li>
  <li>You carry responsibility for other people’s data and money</li>
  <li>You need to comply with rules that do not apply to personal projects (GDPR, accessibility requirements, industry regulations)</li>
  <li>Financial risk: if it breaks, the cost is not just your time</li>
  <li>You have to maintain it, even when you have lost interest</li>
</ul>

<p>A professional software engineer does not just know how to write code. They know how to deal with all of that. How to handle errors that the happy path never reveals. How to design for concurrent users. Where the security boundaries are. What happens at 3 AM when the database fills up and nobody is watching. Not because they are necessarily smarter, but because building for others requires a different kind of thinking.</p>

<h2 id="where-this-is-going">Where This Is Going</h2>

<p>Look at how the construction industry works today. Nobody thinks it is strange that you wallpaper your own living room but call a carpenter for a kitchen renovation. Nobody thinks it is gatekeeping that you need permits and inspections when you build for others. DIY and the professional trades can both exist and are not mutually exclusive.</p>

<p>Software engineering is heading to the same place. People will keep building their own tools with AI, and that is good. But when the stakes go up, when it is someone else’s data, someone else’s money, someone else’s business, they will call in a professional. Not necessarily to write the code, but to assess what was built, identify the risks, and decide whether to build on top of it or start over with a proper foundation. Exactly the way a carpenter assesses a renovation job today.</p>

<p>The evaluation requests I mentioned are the beginning of that pattern. The role of the software engineer is shifting from “the person who writes the code” to “the person who knows whether the code is fit for purpose, and who understands the regulations, compliance, security, and data integrity requirements that come with building for others.” That is a bigger role which requires more expertise, because you need to understand both what was built and what should have been built.</p>

<p>AI will not replace software engineers. It will change what software engineers do.</p>

<p><em>Wondering whether your vibe-coded application is ready for others to use? <a href="#" onclick="task1(); return false;">Get in touch</a> to find out.</em></p>

<h2 id="related-posts">Related Posts</h2>

<ul>
  <li><a href="/ai/development/2026/02/05/vibe-coding-quality-democratisation.html">Vibe Coding: Product Quality and Democratisation</a> – The product quality angle and the bottleneck shift</li>
  <li><a href="/ai/development/testing/2026/06/08/your-ai-tests-are-probably-lying-to-you.html">Your AI Tests Are Probably Lying to You</a> – One example of invisible quality: tests that pass without proving anything</li>
  <li><a href="/ai/llm/development/best-practices/2025/11/14/human-in-the-loop-ai-code-review.html">Human in the Loop: Why Your LLM-Assisted Code Still Needs Human Eyes</a> – Why AI-generated code still needs professional review</li>
</ul>]]></content><author><name>Albert Sikkema</name></author><category term="ai" /><category term="development" /><category term="opinion" /><summary type="html"><![CDATA[AI lets anyone build software, just like DIY lets anyone build a shelf. But building for others has demands that building for yourself does not.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://www.albertsikkema.com/assets/images/why-diy-software-is-great-until-it-is-not-blog.png" /><media:content medium="image" url="https://www.albertsikkema.com/assets/images/why-diy-software-is-great-until-it-is-not-blog.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Code Quality Tools: What the Research Supports</title><link href="https://www.albertsikkema.com/development/quality/research/2026/06/13/what-science-says-about-code-quality-tools.html" rel="alternate" type="text/html" title="Code Quality Tools: What the Research Supports" /><published>2026-06-13T00:00:00+00:00</published><updated>2026-06-13T00:00:00+00:00</updated><id>https://www.albertsikkema.com/development/quality/research/2026/06/13/what-science-says-about-code-quality-tools</id><content type="html" xml:base="https://www.albertsikkema.com/development/quality/research/2026/06/13/what-science-says-about-code-quality-tools.html"><![CDATA[<p>Over the past year I have been working on tooling to improve code quality. That starts with one question first: what constitutes good code? I am a pragmatist, not a purist. Good code in my book is code that the user is happy with (it does what is expected both in use and in function), that is secure and that is easy to maintain.</p>

<p>So how do we measure this? User happiness: just ask for it, and measure some statistics. Security: be diligent and complete (and a bit obsessive), if needed do an audit. But what about maintainability? The goal is to make sure code is maintainable by adhering to a certain minimal standard, enabling developers and LLM’s to work faster, create better code and features more easily and create a better guardrail for code quality assurance.</p>

<p>In order to do that you need to know what works and what does not. What is a marketing hype? What is actually proven to work? Which methods of analysis can be trusted to use as a gate and which are at best an indicator? So I started by implementing what I already knew and some preliminary research which resulted in about 10 methods that I integrated into an mcp server: works fine, and it helps me tremendously to get a quick overview over a codebase, pointing out weaknesses. But at a certain point you keep adding tools and start wondering: do they actually add value to this analysis? The internet is full of tools (like CodeScene and SonarQube) that claim to help, but they all have a slightly different approach.</p>

<p>Given the increased complexity of my tool (and thus the analysis results) and that I plan to also implement the tool I build in CI/CD pipelines, I need to make sure that what I create makes some sense. So before writing any more code for that, I wanted to know: which code quality methods are backed by real evidence? Which ones sound scientific but are not? I get a lot of flagged problems, how to prioritise them? If some of those methods do not have enough grounding, I can delete them. And perhaps it will also tell something about paid options and their credibility and usefulness given costs.</p>

<p>I spent some time over the last months researching this, thinking this would be a solid field where methods would have proper evidence. What I found surprised me: the methods most tools focus on (complexity metrics, code smells, coupling scores) have weaker evidence than you would think.</p>

<p>This post is about the research, not the tools or the implementation. I want to lay out what I found so you can decide for yourself what is worth paying attention to.</p>

<p>Caveat: I am not an expert at code review methods (after all that is why I started this research). Please take it as such: I did my best to not miss any important papers, but if I have, please let me know.</p>

<h2 id="what-is-quality-code-anyway">What Is “Quality Code” Anyway?</h2>

<p>Before diving into which measurement methods work, it helps to define what we are trying to measure, which is quite hard. Fortunately (or not entirely coincidental) my simple definition of what quality code is, resonates in the more official definitions.</p>

<p><a href="https://en.wikipedia.org/wiki/Software_quality">Garvin (1984)</a>, adapted to software by Kitchenham and Pfleeger, identified five competing views of quality:</p>

<ol>
  <li><strong>Transcendental</strong>: “I know it when I see it.” Recognizable but indefinable.</li>
  <li><strong>User-focused</strong>: fitness for purpose. Does it solve the problem? (<a href="https://en.wikipedia.org/wiki/Joseph_M._Juran">Juran</a>’s “fitness for use”)</li>
  <li><strong>Manufacturing</strong>: conformance to requirements and specifications.</li>
  <li><strong>Product-based</strong>: measurable through inherent properties (complexity, coupling, duplication).</li>
  <li><strong>Value-based</strong>: quality relative to cost. Different stakeholders weight it differently.</li>
</ol>

<p>Most arguments about code quality are people talking past each other across these views. A <a href="https://www.sonarsource.com/products/sonarqube/">SonarQube</a> dashboard is view #4. A user who says “it works fine” is view #2. A manager asking “is it worth fixing?” is view #5.</p>

<p>The current international standard, <a href="https://quality.arc42.org/standards/iso-25010">ISO/IEC 25010 (2023)</a>, defines nine product quality characteristics: functional suitability, performance efficiency, compatibility, interaction capability, reliability, security, maintainability, flexibility, and safety. That is a useful checklist, but we can not use it to measure the code directly. What matters for this post is the causal chain it formalizes, originally from <a href="https://www.geeksforgeeks.org/software-engineering/mccalls-quality-model/">McCall (1977)</a> and codified in <a href="https://en.wikipedia.org/wiki/ISO/IEC_9126">ISO 9126 (2001)</a>: internal quality (static code properties like complexity and coupling) <em>influences</em> external quality (runtime behavior like reliability and performance), which influences quality in use (actual user outcomes). Internal quality is necessary but not sufficient. Perfectly structured code that solves the wrong problem is still low quality. But you have to start somewhere, so let’s assume the external quality and usage quality are covered.</p>

<figure>
  <img src="/assets/images/code-quality-research-causal-chain.svg" alt="Causal chain: Internal Quality influences External Quality influences Quality in Use" width="1180" height="380" loading="lazy" style="width:100%;height:auto" />
  <figcaption>The ISO 9126 / McCall causal chain. Code quality tools measure the left box only.</figcaption>
</figure>

<p>The measurement methods I evaluate try to measure internal quality through code properties (a.o. complexity, coupling, duplication, churn patterns) and use those as proxies for maintainability. They do not measure whether users are happy or whether the software solves the right problem. That is an important limitation to keep in mind: even if every metric below worked perfectly, it would only cover one dimension of quality.</p>

<h2 id="the-setup">The Setup</h2>

<p>I looked at 13 different analysis methods that I either already built or considered building: hotspot analysis, code ownership, code smells, cognitive complexity, deep nesting, temporal coupling, duplicate code, coupling metrics, composite health scores, static analysis gates, function length, error handling analysis, and circular dependencies. For each one I collected supporting evidence and counterarguments from peer-reviewed papers.</p>

<p>These methods serve two different purposes. Some try to find specific bugs (static analysis catching a null dereference). Others assess code properties that make maintenance harder or easier (complexity, coupling, smells). Both matter for maintainability and code quality.</p>

<p>So for each method below I look at both: does it help find or predict bugs, and does it give useful insight about maintainability? Then I weigh whether it is worth keeping in a tool.</p>

<h2 id="the-methods">The Methods</h2>

<h3 id="hotspot-analysis-change-frequency-x-complexity">Hotspot Analysis (Change Frequency x Complexity)</h3>

<p>The idea: files that change often AND are complex are where your problems live. <a href="https://dl.acm.org/doi/10.1145/1062455.1062514">Nagappan and Ball (2005)</a> showed that relative churn (normalized by component size) discriminated fault-prone binaries at 89% accuracy on Windows Server 2003 (I had that version running once, that is rather long ago). Raw commit counts are useless; you have to normalize.</p>

<p><a href="https://dl.acm.org/doi/10.1145/3524843.3528091">Tornhill and Borg (2022)</a> found that low-quality code (identified via hotspot analysis) contains 15x more defects and takes 124% longer to work on. This was across 39 proprietary codebases, peer-reviewed at IEEE/ACM TechDebt 2022. Tornhill is the founder of <a href="https://codescene.com/">CodeScene</a>, the tool that implements this approach. The research was done on CodeScene customers.</p>

<p>There is a deeper question though. The “complexity” half of hotspot analysis might just be measuring lines of code. Cyclomatic complexity, introduced by <a href="https://en.wikipedia.org/wiki/Cyclomatic_complexity">McCabe in 1976</a>, counts the number of independent paths through a function: every <code class="language-plaintext highlighter-rouge">if</code>, <code class="language-plaintext highlighter-rouge">for</code>, <code class="language-plaintext highlighter-rouge">while</code>, or <code class="language-plaintext highlighter-rouge">case</code> adds one. The idea is that more paths means harder to test and more likely to contain bugs. It is probably the most widely used complexity metric in the industry. But <a href="https://www.researchgate.net/publication/220204439">Jay et al. (2009)</a> showed that cyclomatic complexity has “absolutely no explanatory power of its own” beyond LOC. The correlation between code complexity and lines of code is so stable across languages and paradigms that they are effectively measuring the same thing.</p>

<p><strong>Verdict: positive.</strong> Be aware the complexity half may just be measuring size, but enough evidence to actually use it.</p>

<h3 id="code-ownership-and-knowledge-distribution">Code Ownership and Knowledge Distribution</h3>

<p><a href="https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/bird2011dtm.pdf">Bird et al. (2011)</a> studied Windows Vista and Windows 7 and found that components with many low-expertise contributors had significantly more defects. Removing ownership features from their prediction model dramatically decreased performance, confirming that ownership is a genuine signal, not a proxy for something else.</p>

<p>This has been replicated across multiple Microsoft products and open-source projects. <a href="https://dl.acm.org/doi/abs/10.1145/2884781.2884852">Thongtanunam et al. (2016)</a> found that code review partially mitigates the ownership effect but does not eliminate it.</p>

<p>The counterargument: shared files may be inherently harder (more complex, more integration-heavy), and that is why many developers touch them. Correlation, not causation. And <a href="https://en.wikipedia.org/wiki/Bus_factor">bus factor</a> (the number of people who would need to disappear before nobody understands a piece of code) is a lagging indicator. By the time you measure that only one person knows how the billing module works, you are already in trouble if that person leaves.</p>

<p><strong>Verdict: positive.</strong> Strongest evidence alongside hotspot analysis. Also a direct maintainability signal: knowing where knowledge is concentrated tells you where onboarding is most critical and where your team is most vulnerable.</p>

<h3 id="cognitive-complexity">Cognitive Complexity</h3>

<p><a href="https://www.sonarsource.com/resources/cognitive-complexity/">SonarSource’s cognitive complexity</a> (the metric behind SonarQube rule S3776) is probably the most widely deployed complexity metric today. And it does correlate with perceived understandability. <a href="https://arxiv.org/pdf/2007.12520">Munoz Baron et al. (2020)</a> confirmed this, and <a href="https://arxiv.org/abs/2303.07722">Lenarduzzi et al. (2023)</a> found it slightly outperforms cyclomatic complexity for readability.</p>

<p>However it has not been validated as a defect predictor. “Harder to understand = more bugs” sounds logical, but nobody has proven this for cognitive complexity specifically. And it inherits the same fundamental critique as cyclomatic complexity: <a href="https://www.semanticscholar.org/paper/A-critique-of-cyclomatic-complexity-as-a-software-Shepperd/a4b522d7d55c0ed38c825c4fb9fe28c14d659c0a">Shepperd (1988)</a> showed that CC is “based upon poor theoretical foundations” and is “no more than a proxy for, and in many cases is outperformed by, lines of code.”</p>

<p>So cognitive complexity tells you something about readability as a formula approximates it, not necessarily how a human developer perceives the code. And it is worth noting the difference: when <a href="https://www.semanticscholar.org/paper/Learning-a-Metric-for-Code-Readability-Buse-Weimer/1a2b8aa0ed7f24ca001508654f506ea010b18a5e">Buse and Weimer (2010)</a> had 120 human annotators rate code readability, that <em>did</em> correlate with fewer defects. But that is human perception, not a formula counting nesting levels and branch points. Cognitive complexity claims to approximate what humans perceive, but the evidence that it actually does so well enough to predict outcomes is surprisingly thin for something so widely adopted.</p>

<p><strong>Verdict: positive.</strong> Weak as a defect predictor, but the maintainability value is clear. A function with a score of 47 is harder for a new team member to understand than one scoring 8.</p>

<h3 id="deep-nesting">Deep Nesting</h3>

<p>Nesting depth (how many levels of <code class="language-plaintext highlighter-rouge">if</code>/<code class="language-plaintext highlighter-rouge">for</code>/<code class="language-plaintext highlighter-rouge">while</code> are stacked inside each other) consistently correlates with fault rates. <a href="https://en.wikipedia.org/wiki/Les_Hatton">Hatton (1997)</a> identified it as one of the most important dimensions that account for defect variability. It is more reliable than function length as a standalone metric.</p>

<p>But it is confounded with function length: deeply nested code is usually in long functions. <a href="https://ieeexplore.ieee.org/document/935855">El Emam et al. (2001)</a> showed that after controlling for size, many metrics lose significance, and nesting may be one of them. There is also no evidence that <em>reducing</em> nesting (via early-return refactoring, for example) improves outcomes. You can scatter logic across more exit points and reduce the measured nesting without making the code any clearer.</p>

<p><strong>Verdict: positive.</strong> The bug-prediction evidence is weakened by El Emam’s finding that nesting may just be a size proxy. It does not predict bugs better than LOC does. But deeply nested code is usually in long functions, and long functions do have more bugs, so the correlation comes along indirectly through size. As a maintainability signal it is more defensible: deeply nested code is genuinely hard to follow regardless of whether it independently predicts defects. Threshold-based (flag anything beyond 3-4 levels) is more useful than continuous measurement. Good supporting signal, not a primary one.</p>

<h3 id="code-smells">Code Smells</h3>

<p>“God class” and “long method” have consistent defect correlation across studies. <a href="https://link.springer.com/article/10.1007/s10664-017-9535-z">Palomba et al. (2018)</a> found that “the great majority of analysed research papers found a positive correlation between code smells and software bugs.”</p>

<p><a href="https://ieeexplore.ieee.org/document/935855">El Emam et al. (2001)</a> placed that into perspective 17 years earlier: out of 24 OO metrics examined, only 4 retained any relationship to faults after controlling for class size. Large classes have more bugs because they have more code. The smell may just be a proxy for size. (more code, more bugs –&gt; makes sense)</p>

<p><a href="https://www.semanticscholar.org/paper/A-survey-on-software-smells-Sharma-Spinellis/6fa6e56c6d57efd5e0fbbe61093c0a0bd32f73cc">Sharma and Spinellis (2018)</a> surveyed the full landscape of smell research, and the evidence per smell type varies a lot. Feature envy (a method that uses another class’s data more than its own) shows weaker and less consistent results. Long parameter list has almost no empirical support. Duplicated code gets the most research attention but its link to defects is debated (see duplicate code below). <a href="https://refactoring.guru/refactoring/smells">Fowler’s original catalog</a> includes more exotic smells (message chains, middle man, speculative generality), but <a href="https://link.springer.com/article/10.1007/s10664-011-9171-y">Khomh et al. (2012)</a> only found that smell-affected classes are more change-prone as a group. That broad finding does not tell you which specific smells drive the effect, and the exotic smells have neither defect evidence nor clear maintainability benefits.</p>

<p>As maintainability indicators, the well-evidenced smells have direct value. A god class is hard to modify regardless of whether it contains bugs. The change-proneness finding from Khomh supports this: smell-affected classes get modified more often, which means developers spend more time on them. That is a maintenance cost even if it never produces a bug.</p>

<p>Smell detectors tend to produce a lot of noise. <a href="https://dl.acm.org/doi/10.1145/1646353.1646374">Bessey et al. (2010)</a>, writing from Coverity’s experience analyzing billions of lines of production code, found that false positives kill adoption. Users prefer fewer true findings over wading through noise.</p>

<p><strong>Verdict: positive for god class and long method,</strong> which have both defect correlation and clear maintainability value. <strong>Negative for exotic smells</strong> (message chains, middle man, speculative generality), which have neither proven defect prediction nor demonstrable maintainability benefits.</p>

<h3 id="temporal-coupling-co-change-analysis">Temporal Coupling (Co-Change Analysis)</h3>

<p>Files that consistently change together may have hidden dependencies that structural analysis misses. <a href="https://dl.acm.org/doi/10.1109/WCRE.2009.5070547">D’Ambros et al. (2009)</a> found that change coupling correlates with defects across three large open-source systems. <a href="https://www.sciencedirect.com/science/article/abs/pii/S0164121214000351">Canfora et al. (2014)</a> found 64-93% of defects in classes with Granger-positive results.</p>

<p>But temporal coupling is noisy. Bulk refactoring, API changes, and rename commits produce false positives. Results are system-dependent: what works for one project may not work for another. It needs aggressive filtering (minimum commit threshold, maximum files per changeset) to be useful.</p>

<p><strong>Verdict: positive.</strong> Unique signal that no other method provides (hidden dependencies), but noisy as a standalone predictor. For maintainability, knowing which files are secretly coupled is valuable for planning refactors. Filter aggressively, present as “hidden dependency finder.”</p>

<h3 id="duplicate-code">Duplicate Code</h3>

<p><a href="https://dl.acm.org/doi/10.1109/ICSE.2009.5070547">Juergens et al. (2009)</a> found that 52% of clones were inconsistently changed and 15% of those inconsistencies caused faults. <a href="https://www.sciencedirect.com/science/article/pii/S0167642310002091">Bettenburg et al. (2012)</a> found only 1-3% of inconsistent changes introduce defects. And <a href="https://link.springer.com/article/10.1007/s10664-011-9195-3">Rahman et al. (2012)</a> found that clones may actually be <em>less</em> defect-prone than non-cloned code, possibly because cloned code tends to be simpler boilerplate.</p>

<p><strong>Verdict: positive.</strong> Not a defect predictor, but genuinely useful for maintainability: when you fix a bug in one copy and forget the other three, that is a real maintenance problem. The value is in reducing the surface area of future changes, not in predicting where bugs are today.</p>

<h3 id="coupling-metrics">Coupling Metrics</h3>

<p>Coupling measures how much one piece of code depends on other pieces. <a href="https://condor.depaul.edu/dmumaugh/OOT/Design-Principles/oodmetrc.pdf">Robert C. Martin</a> proposed a framework that counts incoming dependencies (how many other modules use this one) and outgoing dependencies (how many other modules this one uses) to calculate an “instability” score. Theoretically clean. But <a href="https://www.researchgate.net/publication/31598248_A_Validation_of_Martin's_Metric">Al Dallal (2013)</a> found “a lack of theoretical and empirical evaluation.” The older <a href="https://en.wikipedia.org/wiki/Programming_complexity#Chidamber_and_Kemerer">Chidamber-Kemerer metrics</a> from 1994, particularly their “coupling between objects” metric (simply counting how many other classes a class is connected to), have much better empirical validation. Martin’s metrics add little over that simpler measure. And as we found before, coupling metrics might just be proxies for LOC.</p>

<p><strong>Verdict: positive.</strong> Weak for defect prediction, but hard to dismiss as a maintainability indicator. A module with 30 incoming dependencies is risky to change because any modification can break 30 other places. That is not about bugs, it is about change impact. Simple coupling counts (CBO) are sufficient; the fancier frameworks add little.</p>

<h3 id="composite-health-scores">Composite Health Scores</h3>

<p><a href="https://codescene.com/">CodeScene</a> assigns a 1-10 “code health” score. The general principle (combining multiple metrics beats a single metric) is well supported. And Tornhill and Borg’s 15x defect rate difference is real.</p>

<p>But what does “7.2 health” mean? <a href="https://dl.acm.org/doi/10.5555/580949">Fenton and Pfleeger (1997)</a> point out that combining things measured on different scales (readability, coupling, churn) into one number violates measurement theory. It is like averaging temperature, wind speed, and humidity into a single “weather score.” You get a number, and it looks scientific, but what does it actually tell you? It seems like simplification taken too far.</p>

<p><strong>Verdict: negative.</strong> Neither a reliable defect predictor nor a useful maintainability signal on its own, because it hides which dimensions are actually suffering. The principle of combining signals is sound, but a single number is not actionable. Better to show the individual signals and let the reviewer decide.</p>

<h3 id="static-analysis-quality-gates">Static Analysis Quality Gates</h3>

<p>Static analyzers scan source code without running it, looking for patterns that are known to cause problems. Like using a variable before it is initialized, a null pointer dereference, a SQL query built from unsanitized user input, a resource opened but never closed. Tools like <a href="https://www.sonarsource.com/products/sonarqube/">SonarQube</a>, <a href="https://eslint.org/">ESLint</a> (JavaScript), <a href="https://pylint.readthedocs.io/">Pylint</a> (Python), <a href="https://spotbugs.github.io/">SpotBugs</a> (Java), and <a href="https://codeql.github.com/">CodeQL</a> (GitHub’s query-based static analyzer) all fall in this category. Most work by matching code against a database of known-bad patterns, though some (like CodeQL) do more sophisticated data flow analysis.</p>

<p>77% of projects use at least one according to <a href="https://azaidman.github.io/publications/bellerSANER2016.pdf">Beller et al. (2016)</a>. The economic argument is that it is cheaper to fix early, even with low recall. Many teams use them as quality gates in CI/CD: the build fails if the analyzer finds issues in new code.</p>

<p>But how low is that recall? <a href="https://www.semanticscholar.org/paper/How-Many-of-All-Bugs-Do-We-Find-A-Study-of-Static-Habib-Pradel/72e779539de6a5f9c4d30e512ea9ca4688e02c6a">Habib and Pradel (2018)</a> found that static bug detectors miss the large majority of bugs. Different tools are mostly complementary, each finding different things. A clean report means nothing was found within that tool’s capabilities, not that the code is correct. <a href="https://en.wikipedia.org/wiki/Edsger_W._Dijkstra">Dijkstra (1969)</a> said it decades ago: testing (and analysis) can show the presence of bugs, never their absence.</p>

<p><strong>Verdict: positive.</strong> This is the one method that actually finds specific bugs rather than measuring code properties. Low recall, but what it catches is real. The economic argument holds: even catching 10% of bugs early is cheaper than finding them in production. Just do not pretend a clean report means the code is correct.</p>

<h3 id="function-length">Function Length</h3>

<p>Size correlates with total defect count. But it does NOT reliably predict defect density (bugs per line). And since LOC and cyclomatic complexity are linearly related, measuring both is measuring the same thing twice.</p>

<p><strong>Verdict: negative.</strong> Does not predict defects beyond what LOC already tells you, and does not add maintainability insight beyond what complexity and nesting already capture.</p>

<h3 id="error-handling-analysis">Error Handling Analysis</h3>

<p>Checking for bare except blocks, swallowed errors, ignored error returns (Go’s unchecked <code class="language-plaintext highlighter-rouge">err</code>), <code class="language-plaintext highlighter-rouge">.unwrap()</code> in non-test Rust code, <code class="language-plaintext highlighter-rouge">await</code> without try/catch. These are widely accepted as defects or at minimum bad practice.</p>

<p>There is no large-scale empirical study quantifying defect rates from specific error handling patterns. But the face validity is high: an ignored error return in Go is a bug waiting to happen, and the false positive rate is low when patterns are well-defined. The limitation is that these checks only catch <em>absent</em> error handling, not <em>incorrect</em> handling (wrong recovery action).</p>

<p><strong>Verdict: positive.</strong> Limited academic validation, high practical value. Low false positive rate makes it safe to flag. Both a bug finder (swallowed errors are real bugs) and a maintainability signal (code that ignores errors is fragile).</p>

<h3 id="circular-dependencies">Circular Dependencies</h3>

<p>Circular imports make refactoring, testing, and deployment harder. Detection via DFS on the import graph is deterministic with zero false positives.</p>

<p>But there is limited empirical evidence directly linking cycles to defects. Small cycles may be benign. And this only addresses accidental complexity: fixing a circular dependency does not fix the underlying design problem that caused it.</p>

<p><strong>Verdict: positive.</strong> Zero false positives makes it safe to always flag. Not a defect predictor, but a clear architectural health signal. A codebase with circular dependencies between major modules is harder to maintain and harder to test in isolation.</p>

<h2 id="the-elephant-in-the-room-everything-might-just-be-loc">The Elephant in the Room: Everything Might Just Be LOC</h2>

<p>This is the single most damaging critique in the field, and it applies to almost everything above.</p>

<p><a href="https://ieeexplore.ieee.org/document/935855">El Emam et al. (2001)</a> tested 24 object-oriented metrics and found that most lost their relationship to faults after controlling for class size. <a href="https://www.researchgate.net/publication/220204439">Jay et al. (2009)</a> showed that cyclomatic complexity has “absolutely no explanatory power of its own” beyond lines of code. <a href="https://ieeexplore.ieee.org/document/885631">Graves et al. (2000)</a> found that when LOC is included, complexity metrics add nothing to fault prediction.</p>

<p>The implication is if LOC predicts defects as well as your 47-metric dashboard, your dashboard is an expensive LOC counter with a nicer UI. Any metric that claims to be useful must demonstrate predictive power <em>after controlling for size</em>.</p>

<h2 id="the-counterarguments-that-apply-to-everything">The Counterarguments That Apply to Everything</h2>

<p>Beyond the LOC problem, there are critiques that cut across all methods. I collected them separately because they are worth reading as a group.</p>

<p><strong>Goodhart’s Law.</strong> “When a measure becomes a target, it ceases to be a good measure” (<a href="https://jellyfish.co/blog/goodharts-law-in-software-engineering-and-how-to-avoid-gaming-your-metrics/">Strathern, 1997</a>). Developers split functions to hit a complexity threshold without improving readability. Code coverage targets lead to trivial tests. Any metric feeding into performance reviews will be optimized for the metric, not quality. (this I recognise from daily practice: a gate states that a function is too complex, so I rewrite it to pass the gate, but often that is a tradeoff on other areas)</p>

<p><strong>The Halstead cautionary tale.</strong> <a href="https://en.wikipedia.org/wiki/Halstead_complexity_measures">Halstead (1977)</a> proposed metrics based on operator/operand counts that were widely adopted in tools and standards. Then <a href="https://docs.lib.purdue.edu/cgi/viewcontent.cgi?article=1302&amp;context=cstech">Shen et al. (1983)</a> debunked them: conclusions based on sample sizes less than 10, core assumptions violated, conceptual errors. Metrics can be widely adopted and still be scientifically invalid. Adoption is not validation.</p>

<p><strong>Models do not transfer.</strong> <a href="https://www.sciencedirect.com/science/article/abs/pii/S0164121200000868">Briand et al. (2002)</a> showed that fault-proneness models built on one project often do not transfer to other projects. Default thresholds in any tool are starting points, not universal truths.</p>

<p><strong>False positives kill adoption.</strong> <a href="https://dl.acm.org/doi/10.1145/1646353.1646374">Bessey et al. (2010)</a> again: users prefer fewer true findings over completeness. Their finding was that developers do not care about missed bugs nearly as much as they hate false positives. A false positive wastes your time right now, a false negative is invisible: you cannot be annoyed by a bug the tool never reported. After enough false alarms, developers stop checking findings entirely, and then even the true positives get ignored. Higher precision means some real issues slip through, but a tool nobody trusts has no use at all.</p>

<p><strong>Only accidental complexity is addressable.</strong> <a href="https://worrydream.com/refs/Brooks-NoSilverBullet.pdf">Brooks (1986)</a> distinguished essential complexity (inherent in the problem) from accidental complexity (artifacts of tools and languages). Code quality tools can only address the accidental kind. Design and requirements are where most difficulty lives. No tool changes that.</p>

<p><strong>DeMarco’s nuance.</strong> Tom DeMarco claimed “you cannot control what you cannot measure” in 1982. In <a href="https://ieeexplore.ieee.org/document/5076468/">2009</a> he revisited that position, writing that his earlier work “made the suggestion that metrics are good and therefore more metrics would be better.” He did not reject metrics entirely, but argued that the most important software projects are transformational, and transformation cannot be measured or controlled in advance.</p>

<h2 id="so-what-do-i-do-with-all-this">So What Do I Do With All This?</h2>

<p>Looking at all 13 methods through both lenses (bug finding and maintainability), a pattern emerges. Hotspot analysis and code ownership have the strongest evidence for predicting defects, and they are also directly useful for maintainability. Static analysis actually finds bugs, even if it misses most of them. Everything else (complexity, smells, coupling, duplicates, nesting) is weak as a defect predictor but has varying degrees of value as a maintainability signal.</p>

<p>This is not a complete overview, not by far. There are methods I did not cover (architecture-level metrics, test quality indicators, dependency freshness, among others) and areas where the research is evolving faster than I can read it. I may revisit this in follow-up posts.</p>

<p>Coming back to paying for services that do this for you and whether or not they are adding value: the strongest signals come from git history, which is free in itself. The metrics most paid tools emphasize are the ones with weaker defect-prediction evidence. You are often paying for convenience, visualization, and CI integration rather than for better science. That can be worth it, but go in knowing what the science actually supports. From maintainability standpoint there is a good case to use them, and also it saves you from having to build this yourself (even though that equation dramatically changed with LLMs at your disposal). But it is much more fun to understand what happens and build it yourself (even if you do not use it for you production code, you’ll learn a lot).</p>

<p>I liked this quote from DeMarco in his article <a href="https://www.cs.uni.edu/~wallingf/teaching/172/resources/demarco-on-se.pdf">from 2009</a> so let’s end with that: “Software development is and always will be somewhat experimental. The actual software construction isn’t necessarily experimental, but its conception is. And this is where our focus ought to be. It’s where our focus always ought to have been.”</p>

<p><em>Looking into code quality improvement for your team? <a href="#" onclick="task1(); return false;">Get in touch</a> to compare notes on what works and what does not.</em></p>

<h2 id="sources-mentioned">Sources Mentioned</h2>

<ul>
  <li><a href="https://worrydream.com/refs/Brooks-NoSilverBullet.pdf">No Silver Bullet</a> (Brooks, 1986) – Essential vs accidental complexity</li>
  <li><a href="https://www.semanticscholar.org/paper/A-critique-of-cyclomatic-complexity-as-a-software-Shepperd/a4b522d7d55c0ed38c825c4fb9fe28c14d659c0a">A Critique of Cyclomatic Complexity</a> (Shepperd, 1988) – CC is a LOC proxy with poor foundations</li>
  <li><a href="https://dl.acm.org/doi/10.5555/580949">Software Measurement: A Necessary Scientific Basis</a> (Fenton &amp; Pfleeger, 1997) – Textbook on measurement theory in software</li>
  <li><a href="https://ieeexplore.ieee.org/document/885631">Predicting Fault Incidence</a> (Graves et al., 2000) – LOC makes complexity metrics redundant</li>
  <li><a href="https://ieeexplore.ieee.org/document/935855">The Confounding Effect of Class Size</a> (El Emam et al., 2001) – Most metrics are LOC proxies</li>
  <li><a href="https://www.sciencedirect.com/science/article/abs/pii/S0164121200000868">OO Metrics Don’t Transfer</a> (Briand et al., 2002) – Models built on one project fail on others</li>
  <li><a href="https://docs.lib.purdue.edu/cgi/viewcontent.cgi?article=1302&amp;context=cstech">Software Science Revisited</a> (Shen et al., 1983) – Debunking Halstead metrics</li>
  <li><a href="https://dl.acm.org/doi/10.1145/1062455.1062514">Use of Relative Code Churn Measures</a> (Nagappan &amp; Ball, 2005) – Churn predicts defects at 89% accuracy</li>
  <li><a href="https://ieeexplore.ieee.org/document/5076468/">Software Engineering: An Idea Whose Time Has Come and Gone?</a> (DeMarco, 2009) – “More metrics is better” recanted</li>
  <li><a href="https://dl.acm.org/doi/10.1109/WCRE.2009.5070547">Change Coupling and Defects</a> (D’Ambros et al., 2009) – Change coupling correlates with defects</li>
  <li><a href="https://www.researchgate.net/publication/220204439">Cyclomatic Complexity and Lines of Code</a> (Jay, Hale et al., 2009) – CC has no explanatory power beyond LOC</li>
  <li><a href="https://dl.acm.org/doi/10.1109/ICSE.2009.5070547">Code Clones in the Large</a> (Juergens et al., 2009) – Inconsistent clone changes cause faults</li>
  <li><a href="https://dl.acm.org/doi/10.1145/1646353.1646374">A Few Billion Lines of Code Later</a> (Bessey et al., 2010) – Coverity’s lessons on static analysis adoption</li>
  <li><a href="https://www.semanticscholar.org/paper/Learning-a-Metric-for-Code-Readability-Buse-Weimer/1a2b8aa0ed7f24ca001508654f506ea010b18a5e">Learning a Metric for Code Readability</a> (Buse &amp; Weimer, 2010) – Human readability ratings correlate with defects</li>
  <li><a href="https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/bird2011dtm.pdf">Don’t Touch My Code!</a> (Bird et al., 2011) – Code ownership predicts defects</li>
  <li><a href="https://www.sciencedirect.com/science/article/pii/S0167642310002091">An Empirical Study on Clones</a> (Bettenburg et al., 2012) – Only 1-3% of clone inconsistencies cause defects</li>
  <li><a href="https://link.springer.com/article/10.1007/s10664-011-9171-y">Do Code Smells Hinder Code Changes?</a> (Khomh et al., 2012) – Smell-affected classes are more change-prone</li>
  <li><a href="https://link.springer.com/article/10.1007/s10664-011-9195-3">Cloned Code: Stable Code</a> (Rahman et al., 2012) – Clones may be less defect-prone</li>
  <li><a href="https://www.researchgate.net/publication/31598248_A_Validation_of_Martin's_Metric">A Validation of Martin’s Metric</a> (Al Dallal, 2013) – Martin’s coupling metrics lack validation</li>
  <li><a href="https://www.sciencedirect.com/science/article/abs/pii/S0164121214000351">Defect Prediction with Change Coupling</a> (Canfora et al., 2014) – Granger-based temporal coupling predicts defects</li>
  <li><a href="https://azaidman.github.io/publications/bellerSANER2016.pdf">Analyzing the State of Static Analysis</a> (Beller et al., 2016) – 77% of projects use static analysis</li>
  <li><a href="https://dl.acm.org/doi/abs/10.1145/2884781.2884852">Code Review and Ownership</a> (Thongtanunam et al., 2016) – Code review partially mitigates ownership effect</li>
  <li><a href="https://www.semanticscholar.org/paper/A-survey-on-software-smells-Sharma-Spinellis/6fa6e56c6d57efd5e0fbbe61093c0a0bd32f73cc">A Survey on Software Smells</a> (Sharma &amp; Spinellis, 2018) – Evidence per smell type varies widely</li>
  <li><a href="https://link.springer.com/article/10.1007/s10664-017-9535-z">On the Diffuseness and Impact of Code Smells</a> (Palomba et al., 2018) – Smells correlate with bugs (with size caveats)</li>
  <li><a href="https://www.semanticscholar.org/paper/How-Many-of-All-Bugs-Do-We-Find-A-Study-of-Static-Habib-Pradel/72e779539de6a5f9c4d30e512ea9ca4688e02c6a">How Many of All Bugs Do We Find?</a> (Habib &amp; Pradel, 2018) – Static analyzers miss the large majority of bugs</li>
  <li><a href="https://arxiv.org/pdf/2007.12520">Cognitive Complexity Validation</a> (Munoz Baron et al., 2020) – Validated for readability, not defects</li>
  <li><a href="https://dl.acm.org/doi/10.1145/3524843.3528091">Code Red: The Business Impact of Code Quality</a> (Tornhill &amp; Borg, 2022) – 15x more defects in low-quality code</li>
  <li><a href="https://arxiv.org/abs/2303.07722">Cognitive Complexity vs Cyclomatic Complexity</a> (Lenarduzzi et al., 2023) – Cognitive complexity slightly outperforms CC for readability</li>
</ul>

<h2 id="tools-referenced">Tools Referenced</h2>

<ul>
  <li><a href="https://www.sonarsource.com/products/sonarqube/">SonarQube</a> – Most widely used static analysis platform</li>
  <li><a href="https://codescene.com/">CodeScene</a> – Behavioral code analysis (churn + complexity + ownership)</li>
  <li><a href="https://eslint.org/">ESLint</a> – JavaScript/TypeScript linter</li>
  <li><a href="https://pylint.readthedocs.io/">Pylint</a> – Python static analyzer</li>
  <li><a href="https://spotbugs.github.io/">SpotBugs</a> – Java bug pattern detector</li>
  <li><a href="https://codeql.github.com/">CodeQL</a> – GitHub’s query-based static analyzer</li>
  <li>Adam Tornhill, <a href="https://pragprog.com/titles/atcrime2/your-code-as-a-crime-scene-second-edition/">Your Code as a Crime Scene</a> – The book behind CodeScene’s approach</li>
</ul>

<h2 id="related-posts">Related Posts</h2>

<ul>
  <li><a href="/ai/development/testing/2026/06/08/your-ai-tests-are-probably-lying-to-you.html">Your AI Tests Are Probably Lying to You</a> – Similar theme: green dashboards that hide real problems</li>
  <li><a href="/python/security/privacy/2026/06/01/benchmarking-open-source-pii-detection.html">Benchmarking Open-Source PII Detection</a> – Same approach: benchmark tools, pick the practical winner</li>
  <li><a href="/ai/development/2026/02/05/vibe-coding-quality-democratisation.html">Vibe Coding: Product Quality and Democratisation</a> – Code quality in the AI era</li>
  <li><a href="/ai/llm/development/best-practices/2025/11/14/human-in-the-loop-ai-code-review.html">Human in the Loop: Why Your LLM-Assisted Code Still Needs Human Eyes</a> – Code review quality with AI assistance</li>
  <li><a href="/ai/development/best-practices/2026/03/31/evidence-based-best-practices-ai-guardrails.html">Evidence-Based Best Practices as AI Guardrails (Part 1)</a> – Same evidence-based methodology applied to AI guardrails</li>
</ul>]]></content><author><name>Albert Sikkema</name></author><category term="development" /><category term="quality" /><category term="research" /><summary type="html"><![CDATA[I reviewed research on 13 code quality methods. The ones with strongest evidence are not the ones most tools focus on.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://www.albertsikkema.com/assets/images/code-quality-research-blog.png" /><media:content medium="image" url="https://www.albertsikkema.com/assets/images/code-quality-research-blog.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Your AI Tests Are Probably Lying to You</title><link href="https://www.albertsikkema.com/ai/development/testing/2026/06/08/your-ai-tests-are-probably-lying-to-you.html" rel="alternate" type="text/html" title="Your AI Tests Are Probably Lying to You" /><published>2026-06-08T00:00:00+00:00</published><updated>2026-06-08T00:00:00+00:00</updated><id>https://www.albertsikkema.com/ai/development/testing/2026/06/08/your-ai-tests-are-probably-lying-to-you</id><content type="html" xml:base="https://www.albertsikkema.com/ai/development/testing/2026/06/08/your-ai-tests-are-probably-lying-to-you.html"><![CDATA[<figure>
  <img src="/assets/images/ai-tests-lying.jpg" alt="Sunlight breaking through a dense forest, bright at the center but dark at the edges" width="1920" height="1280" fetchpriority="high" style="width:100%;height:auto" />
  <figcaption>Photo by <a href="https://unsplash.com/@ingmarr">Ingmar</a> on <a href="https://unsplash.com">Unsplash</a></figcaption>
</figure>

<p>As you may have read in an <a href="/ai/llm/development/best-practices/2025/11/14/human-in-the-loop-ai-code-review.html">earlier post</a>, I do not like writing tests. So when an LLM is able to write it for me: great, I let Claude do it. In a review a colleague pointed out that two of my tests did nothing more than verify that mocks return what mocks are told to return. The tests passed, the ci passed but it mainly tested a very obvious thing. And also useless.  I read them myself three or four times and did not see it either.</p>

<p>That was seven months ago and it has been nagging me: how to improve tests while not spending too much time on it. After all I do rely on tests, so they have to have a certain quality. But I do not want to have to write them myself.</p>

<h2 id="some-more-observations">Some more observations</h2>

<p>After that happened I started looking more carefully at AI-generated tests in general. Not just mine, also in code I reviewed for colleagues, and I saw some patterns: Tests that assert a mock returns what you told it to return. Tests that wrap the assertion in a try/except so they pass even when the check fails. Tests with <code class="language-plaintext highlighter-rouge">if status == 200: assert ...</code> that silently skip when the status is something else. Etcetera.</p>

<p>So I went looking at what the academic world has to say about this. Of course this has been researched: <a href="https://dl.acm.org/doi/10.1109/ICSE.2019.00062">“rotten green tests”</a>, coined by Delplanque et al. in 2019. What it means is tests whose assertions never actually execute. The concept of “test smells” goes back even further, to <a href="https://www.researchgate.net/publication/2534882_Refactoring_Test_Code">van Deursen et al. in 2001</a> and <a href="http://xunitpatterns.com/">Meszaros in 2007</a>. So this has been a known problem for 25 years. What is new is that LLMs produce these smells at industrial scale. (Then again the LLM’s training data probably contains at least some of these badly formulated and constructed tests)</p>

<h2 id="llm-is-not-so-good-at-generating-tests">LLM is not so good at generating tests</h2>

<p>I spent some time going through recent studies on LLM-generated tests and the results are not encouraging.</p>

<p><a href="https://arxiv.org/abs/2406.18181">One study</a> tested LLMs on unit test generation across 17 Java projects. Between 34% and 62% of generated tests did not even compile. Of the ones that did compile and run, 75% of the bugs they missed were missed because the tests used boring, normal input values. The LLM picks safe values that exercise the happy path even when the code is broken. If you need to set a value to NaN or pass an empty string or hit an exact boundary to trigger the bug, the LLM will not do that. It generates the test equivalent of “hello world.”</p>

<p>Another <a href="https://arxiv.org/abs/2511.21382">survey of 115 publications</a> on LLM test generation found that the best model they evaluated detected 8 out of 163 real bugs (0.74%).</p>

<p>And <a href="https://www.coderabbit.ai/blog/state-of-ai-vs-human-code-generation-report">CodeRabbit’s analysis</a> of 470 open-source PRs found that AI-co-authored code has 1.7x more issues overall and up to 2.74x more security vulnerabilities than human-written code.</p>

<p>Those numbers are quite high (higher than I recognise from my daily practice, but that has partly to do with the guardrails I use on every LLM interaction).</p>

<p>The real problem is that you can not trust your tests: LLMs often filter out failing tests before showing you results. They throw away the tests that fail (which might be the ones that actually found something) and present you tests that pass. And then teams look at “400 tests, all passing” and feel confident, which is not a confidence that is justified.</p>

<p>One option is to build the tests first (TDD), which works quite well and I do implement that (partly) by defining tests before the actual code is written. However I also found that the building process can only be planned to a certain extent: chaos is always looming (a package that does not work as expected, incomplete view of how the database works etc.) So while I do believe that TDD is a nice idea, I do see more productive progress in a slightly less defined and rigid way of TDD. But however you want to do this: back to the main question: how to make sure that tests are of a proper quality and test things that matter and do not test things that do not matter?</p>

<h2 id="so-i-built-a-thing">So I Built a Thing</h2>

<p>The idea is simple: instead of a generic “review this code” prompt (which misses test smells because it focuses on code quality, not test purpose), give the LLM a specific taxonomy of test defects to hunt for. And crucially, make it read both the test file AND the production code the test covers. You cannot judge whether a test proves anything without seeing what it is supposed to be testing. So after some research I defined several categories that are recognisable to an LLM. See the resources section below if you want to know more.</p>

<p>The <a href="https://github.com/albertsikkema/codebench/tree/main/.claude/skills/review-tests">skill</a> checks for nine categories:</p>

<ul>
  <li><strong>Rotten green tests</strong> – tests that pass without verifying what they claim. Tautologies, conditional assertions, swallowed errors</li>
  <li><strong>Classic test smells</strong> – Assertion Roulette (multiple asserts, no messages), Eager Test (one test doing too much), Mystery Guest (depends on invisible external state)</li>
  <li><strong>Mock abuse</strong> – so many mocks that the test is not testing real behavior anymore. If your test has more mocks than assertions, something is off</li>
  <li><strong>Contract drift</strong> – tests that verify internal method call order instead of observable output. These break on every refactor even when behavior stays the same</li>
  <li><strong>Missing negative paths</strong> – no 401 test, no 403 test, no input validation test. Every auth-protected endpoint needs these</li>
  <li><strong>LLM-generated defects</strong> – hallucinated APIs that do not exist in the codebase, generic-only inputs, type-not-value assertions like <code class="language-plaintext highlighter-rouge">isinstance(result, dict)</code></li>
  <li><strong>Fixture fragility</strong> – tests that depend on execution order, shared mutable state, or fixtures that do too much</li>
  <li><strong>Assertion quality</strong> – loose assertions (<code class="language-plaintext highlighter-rouge">len(x) &gt;= 1</code> when exact count is known), missing assertions, or ten checks on the happy path and one on the error path</li>
  <li><strong>Missing edge cases</strong> – empty collections, boundary values, concurrent access, cleanup after mutation</li>
</ul>

<p>You run it like this:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>/review-tests                          <span class="c"># review changed test files on current branch</span>
/review-tests backend/tests/unit/      <span class="c"># review a specific path</span>
</code></pre></div></div>

<p>Output is as small as possible, with one line per finding:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>tests/test_auth.py:L42: bug: assert mock.return_value == mock.return_value. Tautology. Assert against expected value from spec.
tests/test_api.py:L18: risk: try/except catches AssertionError. Test passes on failure. Remove bare except.
tests/test_search.py:L91: gap: No 401 test for /api/search endpoint. Add unauthenticated request test.
</code></pre></div></div>

<h2 id="the-contradiction-llms-can-create-bad-tests-and-find-them-quite-well">The Contradiction: LLMs can create bad tests and find them quite well</h2>

<p><a href="https://arxiv.org/abs/2506.07594">Santana Jr et al.</a> evaluated LLMs on detecting test smells in Java projects and found detection rates up to 96% for classic smells. LLMs are better at finding test smells than rule-based tools, especially for things like Assertion Roulette and Eager Test that require understanding what the test is trying to do.</p>

<p>The same technology that mass-produces rotten green tests is also the best tool we have for finding them. I noticed this in practice too. Claude is most of the time pretty good at writing tests, but “pretty good” is not good enough when you need tests that catch real bugs. Left to its own devices it gravitates toward happy paths and shallow assertions. But when you give it a specific list of things to look for and force it to read the production code alongside the test, it catches things I miss. Not always, not perfectly, but consistently better than my own manual review.</p>

<h2 id="where-i-use-it">Where I Use It</h2>

<p>I run it when i decide it is useful (large amount of tests) and as part of my automated PR review. The skill is part of <a href="https://github.com/albertsikkema/codebench">codebench</a>, where I collect Claude Code skills, hooks, and configuration. If you use Claude Code for test generation (and you probably do, because writing tests is boring), give it a try.</p>

<p><em>Letting AI write your tests and wondering what slips through? <a href="#" onclick="task1(); return false;">Get in touch</a> to compare notes.</em></p>

<h2 id="resources">Resources</h2>

<h3 id="research">Research</h3>

<ul>
  <li><a href="https://www.researchgate.net/publication/2534882_Refactoring_Test_Code">Refactoring Test Code</a> (van Deursen et al., 2001) – The original test smells paper</li>
  <li><a href="http://xunitpatterns.com/">xUnit Test Patterns</a> (Meszaros, 2007) – 68 patterns for maintainable tests</li>
  <li><a href="https://dl.acm.org/doi/10.1109/ICSE.2019.00062">Rotten Green Tests</a> (Delplanque et al., ICSE 2019) – Tests that pass because their assertions never execute</li>
  <li><a href="https://growing-object-oriented-software.com/">Growing Object-Oriented Software, Guided by Tests</a> (Freeman &amp; Pryce, 2009) – Test behavior, not implementation</li>
  <li><a href="https://dl.acm.org/doi/10.1145/3379597.3387453">Investigating Severity Thresholds for Test Smells</a> (Spadini et al., MSR 2020) – Not all test smell instances are equally harmful</li>
</ul>

<h3 id="llm-test-generation">LLM Test Generation</h3>

<ul>
  <li><a href="https://arxiv.org/abs/2406.18181">On the Evaluation of Large Language Models in Unit Test Generation</a> (Yang et al., 2024) – 34-62% of LLM-generated tests fail to compile; 75% of undetected defects due to missing triggering inputs</li>
  <li><a href="https://arxiv.org/abs/2511.21382">Large Language Models for Unit Test Generation</a> (Alshahwan et al., 2025) – Survey of 115 publications; best model found 8 of 163 bugs</li>
  <li><a href="https://arxiv.org/abs/2506.07594">Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells</a> (Santana Jr et al., 2025) – LLMs detect test smells at up to 96% accuracy</li>
  <li><a href="https://www.coderabbit.ai/blog/state-of-ai-vs-human-code-generation-report">State of AI vs Human Code Generation Report</a> (CodeRabbit, 2025) – AI-co-authored code has 1.7x more issues and 2.74x more security vulnerabilities</li>
</ul>

<h3 id="tools">Tools</h3>

<ul>
  <li><a href="https://github.com/albertsikkema/codebench/tree/main/.claude/skills/review-tests">review-tests skill</a> – The Claude Code skill discussed in this post</li>
  <li><a href="https://github.com/albertsikkema/codebench">codebench</a> – Skills, hooks, and configuration for Claude Code</li>
  <li><a href="https://github.com/affaan-m/everything-claude-code/blob/main/skills/ai-regression-testing/SKILL.md">AI Regression Testing skill</a> – Community skill for AI blind spots in test generation</li>
</ul>

<h3 id="related-posts">Related posts</h3>

<ul>
  <li><a href="/ai/llm/development/best-practices/2025/11/14/human-in-the-loop-ai-code-review.html">Human in the Loop: Why Your LLM-Assisted Code Still Needs Human Eyes</a> – The meta-test story that started this</li>
  <li><a href="/ai/development/automation/2026/05/23/from-prototype-to-production-automated-builds-with-codebuilder.html">From Prototype to Production: Automated Builds</a> – Where automated test review fits in the build pipeline</li>
  <li><a href="/python/security/privacy/2026/06/01/benchmarking-open-source-pii-detection.html">Benchmarking Open-Source PII Detection</a> – Similar approach: benchmark tools, pick the practical winner</li>
</ul>]]></content><author><name>Albert Sikkema</name></author><category term="ai" /><category term="development" /><category term="testing" /><summary type="html"><![CDATA[LLM-generated tests pass without proving anything. I built a review skill backed by test smell research to catch what green CI dashboards hide.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://www.albertsikkema.com/assets/images/your-ai-tests-are-probably-lying-to-you-blog.png" /><media:content medium="image" url="https://www.albertsikkema.com/assets/images/your-ai-tests-are-probably-lying-to-you-blog.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Benchmarking Open-Source PII Detection Across Domains (1/2)</title><link href="https://www.albertsikkema.com/python/security/privacy/2026/06/01/benchmarking-open-source-pii-detection.html" rel="alternate" type="text/html" title="Benchmarking Open-Source PII Detection Across Domains (1/2)" /><published>2026-06-01T00:00:00+00:00</published><updated>2026-06-01T00:00:00+00:00</updated><id>https://www.albertsikkema.com/python/security/privacy/2026/06/01/benchmarking-open-source-pii-detection</id><content type="html" xml:base="https://www.albertsikkema.com/python/security/privacy/2026/06/01/benchmarking-open-source-pii-detection.html"><![CDATA[<figure>
  <img src="/assets/images/pii-detection-benchmark.jpg" alt="Coastal mudflat landscape under dramatic cloudy sky, visibility fading into the distance" width="1920" height="1280" fetchpriority="high" style="width:100%;height:auto" />
  <figcaption>Photo by <a href="https://www.pexels.com/@jornt-hornstra-108388438">Jornt Hornstra</a> on <a href="https://www.pexels.com">Pexels</a></figcaption>
</figure>

<p>About seven months ago I wrote about <a href="/python/security/best-practices/gdpr/2025/11/06/zero-latency-pii-filtering-python-logging.html">zero-latency PII filtering in Python logging</a>: how to sanitize personal data from logs without blocking your main thread. That solved the <em>where</em> and <em>when</em> of filtering. But it implied that detection itself was a solved problem. Throw some regex at it, done. And when handling logs, this is somewhat easier, since you can define certain keys to be filtered etc.</p>

<p>However filtering PII in text is not a solved problem. When I started planning to build PII redaction into a production pipeline that handles legal documents and Dutch-language content (part of a broader push toward <a href="/ai/development/security/2026/04/07/security-by-design-with-project-codeguard.html">security by design</a> in our projects), the question became: which detection tool actually works across all of that?</p>

<p>So first thing to do is get your bearings: find the most common approaches and some datasets and run a benchmark. Five detection approaches, four datasets, two languages (Dutch (since this was the language for the legal documents) and English (testing this will give some more insight and means I can reuse what I learn here)). This post covers what I found. A follow-up (part 2) will take the winners and push them further: more entity types, domain-specific tuning, the stuff you actually need for production.</p>

<p>Warning: this is not a scientific paper. It is a practical comparison I ran to figure out which tool to bet on for our pipeline. The difficulty with comparing these tools is that they are fundamentally different: different architectures, different label schemes, different assumptions about what PII even is. There is no clean apples-to-apples comparison possible. I did my best to make it fair (shared label set, multiple datasets), but four datasets and five models is not enough for strong statistical claims. What it <em>is</em> enough for is picking a direction.</p>

<h2 id="why-comparing-pii-detection-tools-is-hard">Why Comparing PII Detection Tools Is Hard</h2>

<p>Comparing PII detection tools is getting complicated veryfast. Every tool uses different labels, supports different entity types, and has different ideas about what counts as PII.</p>

<p><a href="https://huggingface.co/iiiorg/piiranha-v1-detect-personal-information">Piiranha</a> outputs fine-grained types: GIVENNAME, SURNAME, CITY, STREET, BUILDINGNUM. <a href="https://github.com/microsoft/presidio">Presidio</a> outputs coarse types: PERSON, LOCATION. <a href="https://huggingface.co/urchade/gliner_multi_pii-v1">GLiNER</a> accepts whatever labels you define at inference time. A “person name” in one system is two separate entities in another.</p>

<p>Coverage gaps make it worse. Piiranha detects 17 PII types but has no DATE entity. Presidio has DATE_TIME but cannot detect passwords or usernames. No two tools cover the same set.</p>

<p>And every dataset has its own taxonomy. <a href="https://huggingface.co/datasets/ai4privacy/pii-masking-300k">AI4Privacy</a> uses GIVENNAME/SURNAME. <a href="https://huggingface.co/datasets/gretelai/synthetic_pii_finance_multilingual">Gretel</a> uses “name.” <a href="https://huggingface.co/datasets/nvidia/Nemotron-PII">Nemotron</a> uses first_name/last_name. <a href="https://huggingface.co/datasets/eriktks/conll2002">CoNLL-2002</a> uses PER.</p>

<p>A naive benchmark that ignores all this is misleading. A model scoring F1 0.63 on one dataset and 0.13 on another might not have “collapsed.” It might simply lack entity types that dataset emphasizes (dates accounting for 25% of spans, for example).</p>

<h2 id="three-questions-people-conflate">Three Questions People Conflate</h2>

<p>When evaluating PII detection, three separate questions get mixed together:</p>

<ol>
  <li>How good is the model at detecting entity types it supports?</li>
  <li>How many entity types does it support?</li>
  <li>How well does it generalize to text from different domains?</li>
</ol>

<p>Evaluating all entity types on a single dataset mixes all three. A model might score poorly not because detection is bad, but because it lacks entity types the dataset emphasizes. And a model that scores well might just be evaluated on data similar to its training set.</p>

<p>I wanted to isolate question 1 first. So I defined a shared label set of six entity types that <em>all</em> models can detect: PERSON_NAME, EMAIL, PHONE, LOCATION, CREDITCARD, and IBAN. Every model gets scored only on these. Fair fight.</p>

<p>Question 3 gets answered by testing across four independent datasets. Question 2 (which additional types each model uniquely supports) is left for part 2.</p>

<h2 id="the-models">The Models</h2>

<p>Five detection approaches, covering the major architectural categories:</p>

<p><strong>Regex</strong>: Compiled patterns for email, phone (via Python <code class="language-plaintext highlighter-rouge">phonenumbers</code> for international format parsing), IBAN, credit card (with Luhn validation), and IPv4. Deterministic, no ML.</p>

<p><strong><a href="https://huggingface.co/iiiorg/piiranha-v1-detect-personal-information">Piiranha</a></strong>: 86M-parameter DeBERTa-v3-base, fine-tuned for token classification on 17 PII entity types. Trained on AI4Privacy PII-Masking-300k. Supports English and Dutch.</p>

<p><strong><a href="https://github.com/microsoft/presidio">Presidio</a></strong> (Microsoft, v2.2): Orchestration framework combining spaCy NER (<code class="language-plaintext highlighter-rouge">en_core_web_lg</code> / <code class="language-plaintext highlighter-rouge">nl_core_news_md</code>) with built-in pattern recognizers for structured entities.</p>

<p><strong><a href="https://huggingface.co/urchade/gliner_multi_pii-v1">GLiNER v1</a></strong>: ~209M-parameter zero-shot NER model. Entity types are defined as natural-language descriptions at inference time. You tell it what to look for, and it looks. Schema-agnostic. This is the <a href="https://arxiv.org/abs/2311.08526">GLiNER architecture</a> fine-tuned on synthetic PII data.</p>

<p><strong><a href="https://huggingface.co/fastino/gliner2-base-v1">GLiNER v2</a></strong>: 205M-parameter multi-task model from Fastino Labs. A separate project from v1 with different architecture and training. Included because v2 claims improved multi-task capabilities.</p>

<h2 id="the-datasets">The Datasets</h2>

<p>Four datasets, two languages, different domains:</p>

<table>
  <thead>
    <tr>
      <th>Dataset</th>
      <th>Languages</th>
      <th>Domain</th>
      <th>Why</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><a href="https://huggingface.co/datasets/ai4privacy/pii-masking-300k">AI4Privacy</a></td>
      <td>EN + NL</td>
      <td>Mixed synthetic</td>
      <td>In-distribution baseline for Piiranha (it trained on this)</td>
    </tr>
    <tr>
      <td><a href="https://huggingface.co/datasets/gretelai/synthetic_pii_finance_multilingual">Gretel Finance</a></td>
      <td>EN + NL</td>
      <td>Financial docs</td>
      <td>Out-of-distribution; formal financial text style</td>
    </tr>
    <tr>
      <td><a href="https://huggingface.co/datasets/nvidia/Nemotron-PII">Nemotron-PII</a></td>
      <td>EN only</td>
      <td>50+ industries</td>
      <td>Broadest domain diversity</td>
    </tr>
    <tr>
      <td><a href="https://huggingface.co/datasets/eriktks/conll2002">CoNLL-2002</a></td>
      <td>NL only</td>
      <td>Newspaper</td>
      <td>Gold-standard human annotations; the only non-synthetic dataset</td>
    </tr>
  </tbody>
</table>

<p>The combination is imperfect. No single dataset covers both languages with non-synthetic, human-annotated PII spans. As far as I can tell, such a dataset does not exist publicly.</p>

<h2 id="evaluation">Evaluation</h2>

<p><strong>Accuracy</strong>: Precision, recall, and F1 with relaxed span matching (+/-5 character tolerance). Micro-averaged across all entity types.</p>

<p><strong>Speed</strong>: Median latency per text and throughput in texts/sec. Measured on Apple Silicon M1 Pro CPU.</p>

<p><strong>Reversibility</strong>: Detected spans get replaced with typed placeholders (<code class="language-plaintext highlighter-rouge">[PERSON_NAME_1]</code>, <code class="language-plaintext highlighter-rouge">[EMAIL_1]</code>). Same entity text gets the same placeholder throughout a document. Pass rate = percentage of documents where <code class="language-plaintext highlighter-rouge">restore(redact(text)) == text</code>. This matters: if you cannot perfectly restore the original after redaction, your redaction system is lossy and you will corrupt data.</p>

<h2 id="results">Results</h2>

<h3 id="cross-dataset-accuracy">Cross-Dataset Accuracy</h3>

<table>
  <thead>
    <tr>
      <th>Model</th>
      <th>AI4Privacy</th>
      <th>Gretel</th>
      <th>Nemotron</th>
      <th>CoNLL-2002 (NL)</th>
      <th><strong>AVG</strong></th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><strong>Piiranha</strong></td>
      <td><strong>0.780</strong></td>
      <td>0.169</td>
      <td><strong>0.699</strong></td>
      <td>0.519</td>
      <td><strong>0.542</strong></td>
    </tr>
    <tr>
      <td><strong>GLiNER v1</strong></td>
      <td>0.455</td>
      <td><strong>0.607</strong></td>
      <td>0.484</td>
      <td>0.595</td>
      <td><strong>0.535</strong></td>
    </tr>
    <tr>
      <td>Presidio</td>
      <td>0.359</td>
      <td>0.298</td>
      <td>0.487</td>
      <td><strong>0.780</strong></td>
      <td>0.481</td>
    </tr>
    <tr>
      <td>GLiNER v2</td>
      <td>0.469</td>
      <td>0.373</td>
      <td>0.510</td>
      <td>0.558</td>
      <td>0.478</td>
    </tr>
    <tr>
      <td>Regex</td>
      <td>0.207</td>
      <td>0.179</td>
      <td>0.297</td>
      <td>0.000</td>
      <td>0.171</td>
    </tr>
  </tbody>
</table>

<p>The first thing that jumps out: none of these models are good. The best average F1 across four datasets is 0.542. Commercial systems claim 0.92-0.99. That is a big gap.</p>

<p>The second thing: the top three (Piiranha, GLiNER v1, Presidio) are closer to each other than any of them are to “good enough.” The difference between first and third place is 0.061 F1. A paired t-test across the four datasets shows none of the pairwise differences reach statistical significance (p»0.10, n=4). They are all mediocre, just in different ways. Piiranha swings wildly (0.169 to 0.780). GLiNER v1 is more consistent (0.455 to 0.607). Presidio sits in between.</p>

<h3 id="generalization">Generalization</h3>

<p>How much does each model degrade on unfamiliar text?</p>

<table>
  <thead>
    <tr>
      <th>Model</th>
      <th>In-distribution (AI4Privacy)</th>
      <th>Worst OOD dataset</th>
      <th>Drop</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Piiranha</td>
      <td>0.780</td>
      <td>0.169 (Gretel)</td>
      <td>-78%</td>
    </tr>
    <tr>
      <td>GLiNER v1</td>
      <td>0.455</td>
      <td>0.455 (AI4Privacy)</td>
      <td>0%</td>
    </tr>
    <tr>
      <td>GLiNER v2</td>
      <td>0.469</td>
      <td>0.373 (Gretel)</td>
      <td>-20%</td>
    </tr>
    <tr>
      <td>Presidio</td>
      <td>0.359</td>
      <td>0.298 (Gretel)</td>
      <td>-17%</td>
    </tr>
    <tr>
      <td>Regex</td>
      <td>0.207</td>
      <td>0.000 (CoNLL-2002)</td>
      <td>-100%</td>
    </tr>
  </tbody>
</table>

<p>Piiranha’s 78% drop on financial text looks dramatic, and it is. But GLiNER v1’s “consistent” performance means it never goes above 0.607 either. Presidio’s drop is a moderate 17%. The variance differs (Piiranha std=0.271, GLiNER v1 std=0.077, Presidio std=0.213), but all three models land in the same general territory of “not good enough for production without additional work.” Regex scores zero on CoNLL-2002 because that dataset only has person names and locations.</p>

<h3 id="person-names-the-hard-part">Person Names (The Hard Part)</h3>

<p>Person names are the most common PII entity and the hardest to detect because they are context-dependent (is “Holland” a person or a location?).</p>

<table>
  <thead>
    <tr>
      <th>Model</th>
      <th>AI4Privacy</th>
      <th>Gretel</th>
      <th>Nemotron</th>
      <th>CoNLL-2002</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Piiranha</td>
      <td>0.787</td>
      <td>0.139</td>
      <td>0.749</td>
      <td>0.561</td>
    </tr>
    <tr>
      <td>GLiNER v1</td>
      <td>0.446</td>
      <td>0.564</td>
      <td>0.327</td>
      <td>0.759</td>
    </tr>
    <tr>
      <td>GLiNER v2</td>
      <td>0.435</td>
      <td>0.439</td>
      <td>0.325</td>
      <td>0.641</td>
    </tr>
    <tr>
      <td>Presidio</td>
      <td>0.179</td>
      <td>0.365</td>
      <td>0.360</td>
      <td>0.781</td>
    </tr>
  </tbody>
</table>

<p>Piiranha excels on synthetic data (AI4Privacy, Nemotron) but fails hard on formal financial text and is mediocre on newspaper text. GLiNER v1 and Presidio are more stable but with lower peaks. Regex cannot detect person names at all.</p>

<h3 id="phone-numbers-format-matters">Phone Numbers (Format Matters)</h3>

<table>
  <thead>
    <tr>
      <th>Model</th>
      <th>AI4Privacy</th>
      <th>Gretel</th>
      <th>Nemotron</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>GLiNER v1</td>
      <td>0.874</td>
      <td>0.882</td>
      <td>0.834</td>
    </tr>
    <tr>
      <td>Piiranha</td>
      <td>0.874</td>
      <td>0.707</td>
      <td>0.606</td>
    </tr>
    <tr>
      <td>GLiNER v2</td>
      <td>0.497</td>
      <td>0.377</td>
      <td>0.639</td>
    </tr>
    <tr>
      <td>Presidio</td>
      <td>0.348</td>
      <td>0.390</td>
      <td>0.541</td>
    </tr>
  </tbody>
</table>

<p>GLiNER v1 is the most consistent phone detector across all domains. GLiNER v2 is significantly worse than v1 here. Presidio struggles because spaCy was not designed for this.</p>

<h3 id="email-where-patterns-win">Email (Where Patterns Win)</h3>

<table>
  <thead>
    <tr>
      <th>Model</th>
      <th>AI4Privacy</th>
      <th>Gretel</th>
      <th>Nemotron</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Presidio</td>
      <td>0.960</td>
      <td>0.930</td>
      <td>0.996</td>
    </tr>
    <tr>
      <td>Piiranha</td>
      <td>0.934</td>
      <td>0.553</td>
      <td>0.974</td>
    </tr>
    <tr>
      <td>GLiNER v1</td>
      <td>0.854</td>
      <td>0.576</td>
      <td>0.966</td>
    </tr>
    <tr>
      <td>GLiNER v2</td>
      <td>0.792</td>
      <td>0.568</td>
      <td>0.848</td>
    </tr>
  </tbody>
</table>

<p>Presidio’s pattern-based email recognizer beats all NER models. This is one of the few entity types where pattern matching outperforms learned models. Emails have rigid structure, and a well-written regex does not need to “understand” context. The transformer models still do well on AI4Privacy and Nemotron, but drop on Gretel’s financial documents where email formats are embedded in formal text.</p>

<h3 id="speed">Speed</h3>

<table>
  <thead>
    <tr>
      <th>Model</th>
      <th>Median (ms)</th>
      <th>Texts/sec</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Regex</td>
      <td>0.1</td>
      <td>3,861</td>
    </tr>
    <tr>
      <td>Presidio</td>
      <td>15.1</td>
      <td>63</td>
    </tr>
    <tr>
      <td>Piiranha</td>
      <td>118.5</td>
      <td>8.3</td>
    </tr>
    <tr>
      <td>GLiNER v1</td>
      <td>161.3</td>
      <td>5.9</td>
    </tr>
    <tr>
      <td>GLiNER v2</td>
      <td>197.7</td>
      <td>5.0</td>
    </tr>
  </tbody>
</table>

<p>Presidio is 8-11x faster than transformer models (spaCy’s backbone is small). Among transformers, Piiranha’s 86M parameters make it fastest. All measured on CPU. GPU may change the picture, but the relative speed differences will probably be comparable.</p>

<h3 id="reversibility">Reversibility</h3>

<table>
  <thead>
    <tr>
      <th>Model</th>
      <th>Pass Rate</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Regex</td>
      <td>100%</td>
    </tr>
    <tr>
      <td>Piiranha</td>
      <td>100%</td>
    </tr>
    <tr>
      <td>GLiNER v1</td>
      <td>100%</td>
    </tr>
    <tr>
      <td>GLiNER v2</td>
      <td>100%</td>
    </tr>
    <tr>
      <td>Presidio</td>
      <td>64%</td>
    </tr>
  </tbody>
</table>

<p>Presidio fails on 36% of documents. The cause: its internal tokenization produces character offsets that do not precisely match the original text. This will require further work but can be improved.</p>

<h3 id="does-regex-help">Does Regex Help?</h3>

<p>A reasonable improvement could be to add regex to the mix, so I tested this. Both regex and the model are running on the same text and the results are merged.</p>

<table>
  <thead>
    <tr>
      <th>Model</th>
      <th>Base AVG</th>
      <th>+Regex AVG</th>
      <th>Delta</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Piiranha</td>
      <td>0.542</td>
      <td>0.553</td>
      <td>+0.012</td>
    </tr>
    <tr>
      <td>GLiNER v1</td>
      <td>0.535</td>
      <td>0.536</td>
      <td>+0.001</td>
    </tr>
    <tr>
      <td>Presidio</td>
      <td>0.481</td>
      <td>0.481</td>
      <td>+0.001</td>
    </tr>
  </tbody>
</table>

<p>Conclusion: adding regex to NER is not worth it. The largest gain is +0.012 for Piiranha, and even that is driven entirely by the Gretel dataset where Piiranha is already failing. For GLiNER v1 and Presidio, the improvement is within noise.</p>

<p>Regex actually <em>hurts</em> accuracy on AI4Privacy for all three models (-0.010 to -0.011 F1) because the merge logic occasionally displaces correct NER detections with regex false positives. The model choice matters far more than layering regex on top.</p>

<h2 id="so-which-one-do-you-pick">So Which One Do You Pick?</h2>

<p>If the accuracy differences are not significant and none of the models are production-ready on their own, what <em>does</em> differentiate them?</p>

<table>
  <thead>
    <tr>
      <th>Property</th>
      <th>Piiranha</th>
      <th>GLiNER v1</th>
      <th>Presidio</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Average F1</td>
      <td>0.542</td>
      <td>0.535</td>
      <td>0.481</td>
    </tr>
    <tr>
      <td>Best single dataset</td>
      <td>0.780 (AI4Privacy)</td>
      <td>0.607 (Gretel)</td>
      <td>0.780 (CoNLL-2002)</td>
    </tr>
    <tr>
      <td>Worst single dataset</td>
      <td>0.169 (Gretel)</td>
      <td>0.455 (AI4Privacy)</td>
      <td>0.298 (Gretel)</td>
    </tr>
    <tr>
      <td>Catastrophic failure?</td>
      <td>Yes (Gretel)</td>
      <td>No</td>
      <td>No</td>
    </tr>
    <tr>
      <td>Reversibility</td>
      <td>100%</td>
      <td>100%</td>
      <td>64%</td>
    </tr>
    <tr>
      <td>Speed (median ms)</td>
      <td>119</td>
      <td>161</td>
      <td>15</td>
    </tr>
    <tr>
      <td>Framework/ecosystem</td>
      <td>Model only</td>
      <td>Model only</td>
      <td>Full framework</td>
    </tr>
  </tbody>
</table>

<p>Piiranha and GLiNER v1 are standalone models. They give you detections and that is it. You build the anonymization pipeline, the recognizer registry, the deanonymization logic, the operator pipeline, the multilingual engine configuration, all of it yourself.</p>

<p><a href="https://github.com/microsoft/presidio">Presidio</a> gives you most of that out of the box. Recognizer registry, operator pipeline for anonymization and deanonymization, built-in pattern recognizers for entity types the NER models lack (DATE_TIME, URL), multilingual NLP engine support, and active maintenance by Microsoft. Its accuracy gap to the top two is 0.054-0.061 F1, which is not statistically significant with our sample size.</p>

<p>And to keep in mind: Presidio’s NER backend is replaceable. Its architecture is designed for plugging in custom recognizers. You are not hard bound to just spaCy. You can plug in GLiNER v1, Piiranha, or anything else as a custom recognizer and possible get better detection quality while keeping the framework.</p>

<p>When accuracy is equally mediocre across the board, the framework wins. <strong>Presidio is the logical choice.</strong> Not because it detects PII best (it does not), but because it provides the infrastructure you need regardless of which detection model you use, and you can swap the detection model later without rebuilding everything else. This is a very practical benefit over the others.</p>

<p>The reversibility issue (36% failure rate) is real but can be improved. It is caused by span boundary misalignment in Presidio’s tokenization, not by its detection quality. Post-processing Presidio’s output to align spans with the source text before replacement solves this.</p>

<p>Speed is another Presidio advantage. At 15ms per text it is 8-11x faster than the transformer models. For batch processing this matters less, but for real-time pipelines it is significant.</p>

<h2 id="what-i-learned">What I Learned</h2>

<p>The main takeaway: <strong>none of these tools are good at cross-domain PII detection.</strong> Not the open-source ones I tested, and not the commercial ones either. <a href="https://www.tonic.ai/blog/benchmarking-openai-privacy-filter-pii-detection">Tonic.ai</a> reports F1 0.92-0.99 on their own corpus. <a href="https://docs.aws.amazon.com/ai/responsible-ai/comprehend-detectpii/overview.html">AWS Comprehend</a> claims 0.87-0.91. <a href="https://openai.com/index/introducing-openai-privacy-filter/">OpenAI’s Privacy Filter</a> hits 0.96 on PII-Masking-300k. But these are all in-distribution numbers. When <a href="https://arxiv.org/abs/2604.15776">PIIBench</a> tested eight systems across 10 unified datasets with 48 entity types, the <em>best</em> system scored F1 0.14. OpenAI’s Privacy Filter drops from 0.96 to 0.18-0.65 on Tonic’s out-of-distribution test groups. Cross-domain PII detection is genuinely unsolved.</p>

<p>So in the end the decision stops being about which model detects best, since hey are all mediocre. The decision becomes: which tool is the most practical to build on, the easiest to use, the most actively maintained, and the most complete? That is <a href="https://github.com/microsoft/presidio">Presidio</a>. It gives you the framework, the pipeline, the pattern recognizers, and a swappable NER backend so you can improve detection quality over time without rebuilding everything else.</p>

<p>A few other things worth noting:</p>

<p><strong>In-distribution benchmarks are misleading.</strong> Piiranha’s F1 of 0.78 on its training data drops to 0.17 on financial text. If a model’s benchmark only shows results on familiar data, that number is not useful as a predictor for real-world use.</p>

<p><strong>Zero-shot models generalize better than fine-tuned ones.</strong> GLiNER v1 beats Piiranha on every out-of-distribution dataset. Its schema-agnostic architecture seems to help. This makes it a strong candidate for plugging into Presidio as a custom recognizer.</p>

<p><strong>Hybrid regex+NER is oversold.</strong> Common recommendation in the literature (including <a href="https://arxiv.org/abs/2510.07551">RECAP</a> and other <a href="https://doi.org/10.1038/s41598-025-91846-2">hybrid approaches</a>). Our results show it adds almost nothing (+0.001 F1 for GLiNER v1).</p>

<p><strong>GLiNER v2 is worse than v1 for PII.</strong> The multi-task rewrite hurt NER quality. v2 underperforms v1 on three of four datasets.</p>

<h2 id="what-this-benchmark-does-not-prove">What This Benchmark Does Not Prove</h2>

<p>This is a practical comparison to pick a direction, not a rigorous evaluation. Here is what is wrong with it and why I ran it anyway.</p>

<p><strong>Three of four datasets are synthetic.</strong> AI4Privacy, Gretel, and Nemotron are all machine-generated text. CoNLL-2002 is the only real-world data, and it is Dutch newspaper text from 2000. So when I say “cross-domain,” I mostly mean “across different synthetic data generators.” The gap between synthetic financial text and actual financial documents could be larger than the gap between two synthetic datasets. I used what was publicly available with character-level PII span annotations in both English and Dutch. That combination barely exists. A proper evaluation would include real legal documents, real customer support transcripts, real medical records. Unfortunately they are out of reach to test this on.</p>

<p><strong>AI4Privacy is Piiranha’s training distribution.</strong> I use the validation split, not the training split, so it is not literal data leakage. But validation data from the same generator is not independent. Piiranha’s 0.78 on AI4Privacy is inflated compared to what it would score on truly held-out data, and the “78% drop to Gretel” headline is partly an artifact of that inflated baseline.</p>

<p><strong>The +/-5 character tolerance is a judgment call.</strong> Span matching in NER evaluation is sensitive to boundary definitions. Strict matching (exact character positions) penalizes models for including or excluding a space, a period, or a title like “Mr.” Relaxed matching with a 5-character window allows for these tokenization differences without being so loose that partial detections count as hits. Five characters is roughly one token or one word boundary. I picked it because it felt right for this use case, so it was a choice without real base. Other benchmarks use token-level overlap or IoU thresholds. The choice affects absolute F1 numbers but not the relative ranking between models (all models benefit equally from relaxed matching).</p>

<p><strong>CoNLL-2002 only tests two of six shared entity types.</strong> It has person names and locations but no email, phone, credit card, or IBAN. Presidio’s 0.780 on CoNLL-2002 is driven entirely by spaCy’s person/location NER, not by its pattern recognizers. If CoNLL-2002 had structured PII entities, Presidio’s score there would likely be higher (its email recognizer is the best in the benchmark) and the other models’ scores might shift too. This dataset tests a subset of the shared label set, not the full thing.</p>

<p><strong>No per-entity sample counts.</strong> I do not report how many instances of each entity type appear per dataset. If Gretel has 12 credit card numbers and AI4Privacy has 400, the per-entity tables are not equally powered. A model scoring 0.139 on person names in Gretel might be based on hundreds of spans or dozens.</p>

<p><strong>Six shared entity types is a narrow view.</strong> In practice you need more: dates of birth, social security numbers, IP addresses, Dutch BSN numbers. Piiranha detects 17 types. GLiNER can attempt any type you define (with varying accuracy). The choice depends not just on detection quality but on which entities you need. Part 2 covers this.</p>

<p>All that said: I needed to pick a tool and waiting for a perfect benchmark could take a long time. In the end my question was answered in the sense that there is no ‘best’, just different flavours. And in that light choosing Presidio is the practical choice.</p>

<h2 id="what-comes-next">What Comes Next</h2>

<p>The direction is set: Presidio as the framework, with a better NER model plugged in.</p>

<p>Part 2 will:</p>
<ol>
  <li>Test Presidio with GLiNER v1 and Piiranha as custom recognizers to see how much the detection quality improves while keeping the framework</li>
  <li>Evaluate <a href="https://openai.com/index/introducing-openai-privacy-filter/">OpenAI’s Privacy Filter</a>, which claims F1 0.96 on PII-Masking-300k but drops significantly on out-of-distribution text according to <a href="https://www.tonic.ai/blog/benchmarking-openai-privacy-filter-pii-detection">Tonic.ai’s benchmark</a>. It is relatively new and worth testing on our datasets</li>
  <li>Evaluate on the full set of entity types we actually need for production: person names, emails, phone numbers, physical addresses, dates of birth, social security numbers, IBANs, credit cards, IP addresses, and Dutch-specific identifiers (BSN)</li>
  <li>Fix the reversibility issue (span boundary alignment)</li>
  <li>Determine where we still need to augment with additional detection methods</li>
</ol>

<p><em>Building PII detection into your pipeline and wondering which tool to pick? <a href="#" onclick="task1(); return false;">Get in touch</a> to compare notes.</em></p>

<h2 id="resources">Resources</h2>

<ul>
  <li><a href="https://arxiv.org/abs/2604.15776">PIIBench: A Unified Multi-Source Benchmark for PII Detection</a> - Cross-domain benchmark where all tested systems scored below F1 0.14</li>
  <li><a href="https://www.tonic.ai/blog/benchmarking-openai-privacy-filter-pii-detection">Tonic.ai: Benchmarking OpenAI’s Privacy Filter</a> - Independent cross-domain evaluation showing commercial systems degrade significantly on out-of-distribution text</li>
  <li><a href="https://docs.aws.amazon.com/ai/responsible-ai/comprehend-detectpii/overview.html">AWS Comprehend Detect PII Service Card</a> - AWS’s own accuracy claims per entity type</li>
  <li><a href="https://openai.com/index/introducing-openai-privacy-filter/">OpenAI Privacy Filter</a> - F1 0.96 on PII-Masking-300k, much lower cross-domain</li>
  <li><a href="https://arxiv.org/abs/2311.08526">GLiNER: Generalist Model for Named Entity Recognition</a> (Zaratiana et al., NAACL 2024) - The architecture behind GLiNER v1</li>
  <li><a href="https://doi.org/10.1038/s41598-025-91846-2">Hybrid Rule-based NLP and Machine Learning for PII in Financial Documents</a> (Nature Scientific Reports, 2025)</li>
  <li><a href="https://arxiv.org/abs/2510.07551">RECAP: Hybrid Methods for Multilingual PII Detection</a> (NeurIPS, 2025)</li>
</ul>

<h3 id="models-and-datasets-tested">Models and datasets tested</h3>

<ul>
  <li><a href="https://huggingface.co/iiiorg/piiranha-v1-detect-personal-information">Piiranha v1</a> - DeBERTa-v3-base fine-tuned for PII detection</li>
  <li><a href="https://huggingface.co/urchade/gliner_multi_pii-v1">GLiNER v1 (multi_pii)</a> - Zero-shot NER fine-tuned on synthetic PII data</li>
  <li><a href="https://huggingface.co/fastino/gliner2-base-v1">GLiNER v2 (gliner2-base)</a> - Multi-task NER from Fastino Labs</li>
  <li><a href="https://github.com/microsoft/presidio">Presidio</a> - Microsoft’s PII detection and anonymization framework</li>
  <li><a href="https://huggingface.co/datasets/ai4privacy/pii-masking-300k">AI4Privacy PII-Masking-300k</a> - Mixed synthetic PII dataset (EN + NL)</li>
  <li><a href="https://huggingface.co/datasets/gretelai/synthetic_pii_finance_multilingual">Gretel Synthetic PII Finance</a> - Multilingual financial documents</li>
  <li><a href="https://huggingface.co/datasets/nvidia/Nemotron-PII">Nemotron-PII</a> - NVIDIA’s multi-industry PII dataset</li>
  <li><a href="https://huggingface.co/datasets/eriktks/conll2002">CoNLL-2002</a> - Dutch/Spanish NER with human annotations</li>
</ul>]]></content><author><name>Albert Sikkema</name></author><category term="python" /><category term="security" /><category term="privacy" /><summary type="html"><![CDATA[I tested five PII detection tools across four datasets. None are good. When accuracy is equally mediocre, the framework matters more than the model.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://www.albertsikkema.com/assets/images/benchmarking-open-source-pii-detection-blog.png" /><media:content medium="image" url="https://www.albertsikkema.com/assets/images/benchmarking-open-source-pii-detection-blog.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">From Prototype to Production: Automated Builds</title><link href="https://www.albertsikkema.com/ai/development/automation/2026/05/23/from-prototype-to-production-automated-builds-with-codebuilder.html" rel="alternate" type="text/html" title="From Prototype to Production: Automated Builds" /><published>2026-05-23T00:00:00+00:00</published><updated>2026-05-23T00:00:00+00:00</updated><id>https://www.albertsikkema.com/ai/development/automation/2026/05/23/from-prototype-to-production-automated-builds-with-codebuilder</id><content type="html" xml:base="https://www.albertsikkema.com/ai/development/automation/2026/05/23/from-prototype-to-production-automated-builds-with-codebuilder.html"><![CDATA[<p>Last month I wrote about <a href="/ai/development/automation/2026/04/17/automated-builds-cost-fatigue-ceiling.html">where automated builds hit their ceiling</a>: costs, reviewer fatigue, and models that are not smart enough for full autonomy. That post ended with “I am more or less at a standstill when it comes to further improvement.” What I meant with that is the quality of the code: a ceiling seems to be reached with that. However that does not mean it is not good enough to actually use for making real products. I have been working on automated workflows for code building for quite some time, and the next step was presenting the whole thing to my colleagues and hearing “can we use this?”</p>

<h2 id="the-pitch">The Pitch</h2>

<p>I have been iterating on automated builds for over a year now. Different tools, different approaches, lots of starting over. This evolved into a Go server with a React frontend that could orchestrate <a href="https://www.anthropic.com/product/claude-code">Claude Code</a> containers to plan, build, and review code autonomously. It ran on my local server and it worked. I tested it on my own and a few work repos and the output was good: decent quality code and a lot of time saved.</p>

<p>When I showed this to the team, the reaction was “how soon can we plug this into our workflow?” The appeal was straightforward: we can do more with the same team, product owners and domain experts can contribute ideas directly, and we are less dependent on developers for every small change.</p>

<h2 id="used-to-jira">Used to Jira</h2>

<p>The prototype proved the concept but it was not a fit for company use. It replaced too much: it had its own project management UI, backlog and task board. Nice for my own use, but we already have <a href="https://www.atlassian.com/software/jira">Jira</a>. We are used to it (not addicted to it), so it seemed a good choice not to change everything at once: do not replace Jira. Everything still works without the LLM layer. Issues get created in Jira the normal way. Developers can pick them up the normal way. The automated builder is an addition, not a replacement.</p>

<p>This matters because nobody had to change how they work on day one. A product owner creates a Jira issue, adds it to the sprint, and if they put it in a specific status column with a specific label the builder picks it up. If the ticket does not have the label, a human developer picks it up instead. Same board, same workflow, one extra option.</p>

<p>I can imagine we move away from Jira eventually. But now is not the right moment. We need to gain experience with this step first, and forcing a tool change at the same time would muddy the results.</p>

<h2 id="how-it-works">How It Works</h2>

<p>The system is called Codebuilder for internal use. It polls <a href="https://www.atlassian.com/software/jira">Jira</a> for issues in a trigger status, spawns <a href="https://www.docker.com/">Docker</a> containers running Claude Code, and reports results back as comments, status transitions, labels, and pull requests.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Jira (issue moves to trigger status)
  |
  v
Poller --&gt; Engine --&gt; Docker Container
                        |
                        v
                    Claude Code + build script
                        |
                        +-- calls back to Codebuilder API
                        +-- uses MCP server for project knowledge + Jira
                        +-- creates draft PR on GitHub
                        |
                        v
                    Container exits --&gt; Engine completes job
                        |
                        v
                    Jira comment + status transition
</code></pre></div></div>

<p>The worker inside the container does planning and building in one pass. It has access to project knowledge (specifications, requirements, architectural decisions) through an <a href="https://modelcontextprotocol.io/specification/2025-03-26/basic/transports#streamable-http">MCP</a> server, and it can read and update the Jira issue it is working on. When it is done, it creates a draft PR. A developer reviews, possibly makes small changes, and merges. The merged PR auto-transitions the Jira issue to done.</p>

<p>Half an hour from “issue in queue” to “PR ready for review.”</p>

<figure>
  <img src="/assets/images/codebuilder_1.jpg" alt="Codebuilder jobs dashboard showing a list of build and plan jobs with status, PR links, and timestamps" width="1920" height="480" fetchpriority="high" style="width:100%;height:auto" />
  <figcaption>The jobs dashboard. Each row is one Jira issue that went through the pipeline.</figcaption>
</figure>

<h2 id="what-it-looks-like">What It Looks Like</h2>

<p>In essence there are five steps that run in a single docker container, every step is a new claude code session:</p>

<ul>
  <li>research</li>
  <li>plan</li>
  <li>build</li>
  <li>review</li>
  <li>fix</li>
</ul>

<figure>
  <img src="/assets/images/codebuilder_2.jpg" alt="Codebuilder job detail showing completed Research, Plan, Build, Review, Triage, and Fix steps with logs, token usage, cost, and runtime" width="1920" height="960" loading="lazy" style="width:100%;height:auto" />
  <figcaption>A single build run: all steps green, 31 minutes, $8.81 in tokens, PR created. The logs show Claude reasoning about what to stage and what to skip.</figcaption>
</figure>

<p>The reason why everything runs in docker containers is mainly security: this isolates claude code from the server, making sure the LLM (which runs unsupervised), has no access to the server. Also we make sure no secrets are available (like a .env file), to minimise risks. This does mean that small errors can be part of the build, but this is acceptable over allowing a very knowledgeable but non-responsible entity roaming around your files.</p>

<p>What we added: <a href="https://learn.microsoft.com/en-us/entra/identity-platform/v2-protocols-oidc">Azure AD</a> single sign-on, <a href="https://axiom.co/">Axiom</a> for structured logging, Telegram notifications, CSRF protection, rate limiting, and a CLI tool for developers to review and merge automated PRs from their terminal that integrates with the mcp server to make sure github and jira get updated through the cli as well.</p>

<p>Another important one: product owners and others can interact with Jira through Claude Code directly in the browser. We already have the setup to run Claude Code on the server, so by exposing that through a web interface someone can use the full setup without leaving their browser.</p>

<p>This helps tremendously with refining issues and keeping the project up to date. Nothing nicer than asking “what small issues and bugs are there on the backlog that we could combine in one run?” without having to go through a long list yourself.</p>

<h2 id="the-bigger-picture-not-just-for-developers">The Bigger Picture: Not Just for Developers</h2>

<p>The part I am most excited about is what this enables for people who are not developers.</p>

<p>The idea: a product owner or domain expert fills out a form describing what they want. That goes into a refinement queue. An LLM (with human oversight) helps refine the idea into a well-scoped issue with acceptance criteria. The refined issue either goes to the automated build queue or gets assigned to a human developer, depending on complexity and judgment.</p>

<p>This is not “non-technical people writing code.” It is non-technical people contributing meaningful input to the development process without needing a developer to translate their ideas into tickets. The LLM handles the translation, a developer or product owner reviews it, and the builder does the implementation.</p>

<h2 id="what-worries-me">What Worries Me</h2>

<p>Two things, and they are the same ones from my <a href="/ai/development/automation/2026/04/17/automated-builds-cost-fatigue-ceiling.html">previous post</a>.</p>

<p><strong>Token costs.</strong> The quality is high because we run rigorous planning and review cycles, both LLM and human. But the planning and review steps require the best models to maintain a minimum quality bar. Cheaper models produce reviews that miss subtle issues or flag correct code as broken, and bad automated reviews are worse than no reviews because they create false confidence. Token prices keep shifting, and not always downward. That uncertainty makes it hard to forecast costs for the next quarter. In the end it is still cheaper than a purely human review, but the gap is not as large as it used to be.</p>

<p><strong>Developer fatigue.</strong> Reviewing LLM-generated code is fundamentally different from writing code yourself. You lose the mental map you build when you write the code. You are reading someone else’s work in a codebase you did not shape, and the entity that wrote it does not learn from your feedback across sessions. After a few hours your attention drops. After a few days you start wondering if this is what the job looks like from now on. I wrote about this before and it is still true, just more visible now that multiple people experience it instead of only me.</p>

<p>Our current answer: all PRs get an automated LLM review as part of the workflow. The developer decides whether to also ask a colleague for a human review. The guideline is to always request one for database changes, large frontend or backend changes, and anything security-related. At first this seems a bit uneasy, not having a second reviewer. However, the automated review is so thorough that it catches more than most (if not all) humans would.</p>

<h2 id="where-we-are-now">Where We Are Now</h2>

<p>We are in the “use it carefully and learn” phase. It works, it is fast, the code quality is good, and it is cheaper than a developer hour. A year ago I started stitching together shell scripts to run Claude Code in sequence. Now a product owner can drop an idea into Jira and get a reviewable PR back before lunch.</p>

<p>There is a lot more to cover: how we handle failures, what the prompts look like, how we scope issues for the builder, token cost breakdowns, and what happens when the model gets it wrong. That will be a follow-up post once we have more production hours under our belt.</p>

<hr />

<p><em>Want something like this for your company? I can build it, and it is a lot of fun to do. <a href="#" onclick="task1(); return false;">Get in touch</a> and we will figure out what fits your workflow.</em></p>

<h2 id="resources">Resources</h2>

<ul>
  <li><a href="https://www.anthropic.com/product/claude-code">Claude Code</a> - Anthropic’s agentic coding CLI</li>
  <li><a href="https://modelcontextprotocol.io/">Model Context Protocol</a> - The protocol Codebuilder uses to expose project knowledge to workers</li>
  <li><a href="https://axiom.co/">Axiom</a> - Structured logging and observability</li>
  <li><a href="https://learn.microsoft.com/en-us/entra/identity-platform/v2-protocols-oidc">Azure AD OIDC</a> - Authentication setup
    <h3 id="related-posts">Related posts</h3>
  </li>
</ul>

<p>These cover the individual steps and lessons behind Codebuilder in more detail:</p>

<ul>
  <li><a href="/AI/LLM/development/productivity/2025/11/21/orchestrator-automating-claude-code-workflows.html">The Orchestrator: Automating Full Claude Code Workflows</a> - The earlier orchestration approach: research, plan, build, review chained together</li>
  <li><a href="/ai/development/operations/2026/04/16/when-llms-actually-deliver.html">When LLMs Actually Deliver</a> - The tooling that makes LLM builds produce usable output</li>
  <li><a href="/ai/llm/development/best-practices/2025/11/14/human-in-the-loop-ai-code-review.html">Human in the Loop: Why Your LLM-Assisted Code Still Needs Human Eyes</a> - Why automated review alone is not enough</li>
  <li><a href="/ai/development/automation/2026/04/17/automated-builds-cost-fatigue-ceiling.html">Fully Automated LLM Builds: Where It Actually Stops</a> - Costs, fatigue, and model ceilings at scale</li>
  <li><a href="/ai/development/tools/2026/04/23/smaller-context-window-better-claude-code.html">Why I Shrunk Claude Code’s Context Window Back to 200k</a> - Context management lessons that feed into how workers are configured</li>
  <li><a href="/ai/tools/productivity/2025/10/14/supercharge-claude-code-with-custom-configuration.html">Supercharge Claude Code with Custom Configuration</a> - The agent/skill/hook setup that Codebuilder workers inherit</li>
</ul>]]></content><author><name>Albert Sikkema</name></author><category term="ai" /><category term="development" /><category term="automation" /><summary type="html"><![CDATA[How a personal prototype for automated LLM builds became a company-wide platform on top of Jira. Architecture decisions, reasoning, and early results.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://www.albertsikkema.com/assets/images/codebuilder_1.jpg" /><media:content medium="image" url="https://www.albertsikkema.com/assets/images/codebuilder_1.jpg" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Why I Shrunk Claude Code’s Context Window Back to 200k</title><link href="https://www.albertsikkema.com/ai/development/tools/2026/04/23/smaller-context-window-better-claude-code.html" rel="alternate" type="text/html" title="Why I Shrunk Claude Code’s Context Window Back to 200k" /><published>2026-04-23T00:00:00+00:00</published><updated>2026-04-23T00:00:00+00:00</updated><id>https://www.albertsikkema.com/ai/development/tools/2026/04/23/smaller-context-window-better-claude-code</id><content type="html" xml:base="https://www.albertsikkema.com/ai/development/tools/2026/04/23/smaller-context-window-better-claude-code.html"><![CDATA[<figure>
  <img src="/assets/images/context-window-rain-glass.jpg" alt="Rain-covered glass with blurred warm lights behind, signal obscured by noise" width="1920" height="1078" fetchpriority="high" style="width:100%;height:auto" />
  <figcaption>Signal obscured by noise. Photo by <a href="https://unsplash.com/@c_g_">c g</a> on <a href="https://unsplash.com">Unsplash</a></figcaption>
</figure>

<p>This morning I watched a video on context window management in Claude Code as part of my daily “keep up with what is happening in the LLM space” routine. Good content, solid diagnosis of the problem. But everything in it was about manual interventions: trigger compaction at the right moment, use structured handoffs, rewind instead of correcting. All valid techniques. But there is an issue that is buried deeper and there are two settings, buried in the documentation, that may solve most of this.</p>

<h2 id="the-problem-with-more-room">The Problem With More Room</h2>

<p>Since Opus 4.6 it ships with a default <a href="https://claude.com/blog/1m-context-ga">1M token context window</a>. Five times the 200k window of their predecessors. Sounds like a pure upgrade. Great!</p>

<p>It is not. The single most important thing for working with LLMs is context management: keep it as small as possible with as relevant info as possible and nothing more than that. Anthropic’s <a href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents">engineering team</a> says you should be “striving for the minimal set of information that fully outlines your expected behavior.” Their <a href="https://platform.claude.com/docs/en/build-with-claude/context-windows">documentation</a> acknowledges that “as token count grows, accuracy and recall degrade, a phenomenon known as context rot.” The root cause is the n-squared attention mechanism: double the context, quadruple the number of pairwise relationships the model has to track.</p>

<p>I am not alone in experiencing this. Some quick research finds similar sentiment: <a href="https://simonwillison.net/2025/Jan/26/paul-gauthier/">Paul Gauthier</a>, the creator of Aider, found that “every model seems to get confused when you feed them more than ~25-30k tokens.” He calls it the number one problem his users report. <a href="https://blog.jetbrains.com/research/2025/12/efficient-context-management/">JetBrains Research</a> tested observation masking (hiding old tool outputs) and found a 52% cost reduction while <em>boosting</em> solve rates by 2.6%. Less context, better results. The <a href="https://eval.16x.engineer/blog/llm-context-management-guide">NoLiMa benchmark</a> found that 11 of 12 tested models dropped below 50% of their short-context performance at just 32k tokens. Not 200k. Not 1M. 32 thousand.</p>

<p>In practice this means: more hallucinations, forgotten instructions, goal drift, inconsistent decisions. Not at 900k tokens. Much, much earlier.</p>

<h2 id="what-i-keep-seeing">What I Keep Seeing</h2>

<p>A lot of the advice I come across focuses on manual interventions. Trigger compaction yourself at the right moment. Use a new session or <code class="language-plaintext highlighter-rouge">/clear</code> when switching tasks. Save state to a JSON file before clearing. Ask Claude for periodic summaries. Use sub-agents to keep intermediate work out of your main context.</p>

<p>These are all valid. I use sub-agents heavily (they get their own fresh context window, which is <a href="https://www.morphllm.com/context-rot">the single most effective architectural pattern</a> for avoiding context rot) and <code class="language-plaintext highlighter-rouge">/clear</code> between unrelated tasks. But manual interventions are workarounds for a window that is too large, not fixes for the underlying problem. They require you to watch your context usage while trying to get work done. That is overhead the tooling should handle.</p>

<p><a href="https://tessl.io/blog/amp-retires-compaction-for-a-cleaner-handoff-in-the-coding-agent-context-race">Amp</a> went further: they dropped compaction entirely and designed around short threads with clean handoffs. Their senior engineer Dan Mac put it bluntly: “You should basically never use compaction.”.</p>

<h2 id="the-simpler-fix">The Simpler Fix</h2>

<p>Two <a href="https://code.claude.com/docs/en/env-vars">environment variables</a> solve this without ongoing attention:</p>

<p><strong><code class="language-plaintext highlighter-rouge">CLAUDE_CODE_DISABLE_1M_CONTEXT</code></strong> set to <code class="language-plaintext highlighter-rouge">1</code> caps the context window back to 200k tokens. It removes the 1M model variants from the <a href="https://code.claude.com/docs/en/model-config#extended-context">model picker</a> entirely.</p>

<p><strong><code class="language-plaintext highlighter-rouge">CLAUDE_AUTOCOMPACT_PCT_OVERRIDE</code></strong> set to a value between 1-100 controls when auto-compaction triggers, as a percentage of context capacity. The <a href="https://code.claude.com/docs/en/how-claude-code-works">default is around 95%</a>, which means on a 1M window, compaction does not kick in until you are at 950k tokens. That is way past the point where quality has degraded.</p>

<p>Set both in your project or user settings:</p>

<div class="language-json highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="p">{</span><span class="w">
  </span><span class="nl">"env"</span><span class="p">:</span><span class="w"> </span><span class="p">{</span><span class="w">
    </span><span class="nl">"CLAUDE_CODE_DISABLE_1M_CONTEXT"</span><span class="p">:</span><span class="w"> </span><span class="s2">"1"</span><span class="p">,</span><span class="w">
    </span><span class="nl">"CLAUDE_AUTOCOMPACT_PCT_OVERRIDE"</span><span class="p">:</span><span class="w"> </span><span class="s2">"70"</span><span class="w">
  </span><span class="p">}</span><span class="w">
</span><span class="p">}</span><span class="w">
</span></code></pre></div></div>

<p>200k window at 70% threshold means compaction triggers around 140k tokens. Well before quality drops off. Compaction runs more frequently but with less context to summarize, which means better summaries. <a href="https://x.com/karpathy/status/1937902205765607626">Andrej Karpathy</a> described context engineering as “filling the context window with just the right information for the next step.” Too much, and “performance might come down.” A constrained window forces that discipline automatically.</p>

<h2 id="the-cost-angle">The Cost Angle</h2>

<p>Every turn in Claude Code sends the full conversation context to Anthropic’s servers. If your context window is sitting at 600k tokens of accumulated tool output, file reads, and old conversation, all of that gets re-sent and re-billed on the next message.</p>

<p>With <a href="https://platform.claude.com/docs/en/about-claude/pricing">Opus 4.6 at $5 per million input tokens</a>, a 600k context costs a lot per turn just for input (not exactly 3 dollar because there is also caching going on). A 140k context (right before compaction) theoretically costs $0.70. Over a long session with dozens of turns, that difference adds up. One user on <a href="https://news.ycombinator.com/item?id=47580395">Hacker News</a> described burning through $100 in credit when Opus 4.6 “got stuck in a bullshit reasoning loop.” So a smaller window is better for quality and is cheaper.</p>

<h2 id="compaction-is-not-a-safety-net">Compaction Is Not a Safety Net</h2>

<p>The standard advice is to rely on compaction to keep your context clean. And compaction works, sort of and sometimes. It summarizes the conversation to free up space. But it is lossy. The model decides what matters and what gets dropped, and its judgment is not always yours. Often I feel like I have to start over again with the nuances of the problem we were working on.</p>

<p>Another problem is what happens <em>after</em> compaction. Yesterday I was in a session, working on changes across a repo. Auto-compaction kicked in mid-task. First thing Claude did after compacting: committed everything. Without being asked. It lost enough context to forget that I had not asked for a commit, saw uncommitted files and went ahead.</p>

<p><a href="https://claude.com/blog/using-claude-code-session-management-and-1m-context">Thariq Shihipar</a> from the Claude Code team recommends compacting proactively at 50-60% capacity instead of waiting for auto-compaction. Good advice. But if you constrain your window to 200k and set the threshold to 70%, you get roughly the same effect automatically. No need to watch your token count and manually trigger <code class="language-plaintext highlighter-rouge">/compact</code> at the right moment.</p>

<p>There is another upside to the smaller window: compaction itself gets better. If compaction triggers at 70% and reduces back to roughly 30%, the 1M window has to summarize away 400,000 tokens of conversation. The 200k window only discards 80,000. Five times less information to lose. Smaller contexts lead to more accurate summaries. (For the information theory purists out there: I am aware that it is more complicated than this, please forgive my shortcuts)</p>

<h2 id="fresh-sessions-not-long-ones">Fresh Sessions, Not Long Ones</h2>

<p>The env vars help within a session. But the bigger win is avoiding compaction by not having long sessions in the first place.</p>

<p>My workflow separates every phase into its own session. Research runs in one session, planning in another, building in a third. Never chained together in the same conversation. I <a href="/AI/LLM/development/productivity/2025/11/21/orchestrator-automating-claude-code-workflows.html">built an orchestrator</a> that does this automatically: each step launches a separate Claude Code instance. The whole rationale is context isolation. Each phase starts clean with only the information needed to start that session.</p>

<p>This is the same principle behind sub-agents, just at a larger scale. When Claude Code spawns a sub-agent, that agent gets its own fresh context window. All intermediate work (file reads, grep output, failed attempts) stays in the sub-agent’s context. Only the final result comes back. <a href="https://www.morphllm.com/context-rot">Morph’s research</a> found a 90% performance gain using sub-agent architecture over single-agent. The reason is straightforward: every file read and tool call that stays out of your main context is noise that never competes for attention.</p>

<h2 id="the-counterintuitive-takeaway">The Counterintuitive Takeaway</h2>

<p>The 1M context window is a capacity increase, not a quality increase. More room means more space for noise, higher bills, and worse output once you cross the degradation threshold. Steve Smith <a href="https://blog.nimblepros.com/blogs/context-windows-wont-grow-forever/">calls it</a> “a huge junk drawer.” Glen Rhodes <a href="https://glenrhodes.com/context-window-management-treating-llm-context-as-working-memory-not-unlimited-storage/">describes context</a> as working memory, not storage, and argues you should treat it like RAM on a constrained system: deliberate about what gets loaded, suspicious of anything that lingers.</p>

<p>The best results I get from Claude Code come from keeping the window small, avoid compacting at all, and never letting one phase pollute the next. Two environment variables and a habit of starting fresh. That is the whole trick.</p>

<hr />

<p><em>Running into context issues or managing your own Claude Code setup? <a href="#" onclick="task1(); return false;">Get in touch</a> to compare notes.</em></p>

<h2 id="resources">Resources</h2>

<h3 id="claude-code-documentation">Claude Code documentation</h3>

<ul>
  <li><a href="https://claude.com/blog/using-claude-code-session-management-and-1m-context">Using Claude Code: session management and 1M context</a> – Thariq Shihipar’s practical guide to context management</li>
  <li><a href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents">Effective context engineering for AI agents</a> – Anthropic’s engineering blog on context rot and mitigation</li>
  <li><a href="https://code.claude.com/docs/en/env-vars">Claude Code environment variables</a> – Official docs including the two env vars discussed here</li>
  <li><a href="https://code.claude.com/docs/en/context-window">Explore the context window</a> – Interactive simulation of how context fills during a session</li>
  <li><a href="https://platform.claude.com/docs/en/build-with-claude/context-windows">Context windows API docs</a> – Model-by-model context sizes and Anthropic’s acknowledgment of context rot</li>
</ul>

<h3 id="research-and-analysis">Research and analysis</h3>

<ul>
  <li><a href="https://arxiv.org/abs/2307.03172">Lost in the Middle (Liu et al., 2023)</a> – The foundational research on performance degradation in long contexts</li>
  <li><a href="https://www.trychroma.com/research/context-rot">Context Rot research by Chroma</a> – Evaluation of 18 LLMs showing universal degradation with input length</li>
  <li><a href="https://blog.jetbrains.com/research/2025/12/efficient-context-management/">JetBrains: Smarter Context Management for Agents</a> – Observation masking: less context, better results</li>
  <li><a href="https://www.morphllm.com/context-rot">Morph: Context Rot complete guide</a> – Sub-agent architecture and agent-specific context data</li>
  <li><a href="https://gist.github.com/badlogic/cd2ef65b0697c4dbe2d13fbecb0a0a5f">Compaction research across coding tools</a> – Claude Code, Codex CLI, OpenCode, Amp compared</li>
</ul>

<h3 id="developer-perspectives">Developer perspectives</h3>

<ul>
  <li><a href="https://simonwillison.net/2025/Jan/26/paul-gauthier/">Paul Gauthier on practical context limits</a> – Aider creator: models get confused above 25-30k tokens</li>
  <li><a href="https://x.com/karpathy/status/1937902205765607626">Karpathy on context engineering</a> – “Too much or too irrelevant, and performance might come down”</li>
  <li><a href="https://tessl.io/blog/amp-retires-compaction-for-a-cleaner-handoff-in-the-coding-agent-context-race">Amp drops compaction for handoff</a> – Why one coding tool designed around short threads</li>
  <li><a href="https://blog.nimblepros.com/blogs/context-windows-wont-grow-forever/">Why Context Windows Won’t Keep Growing Forever</a> – Steve Smith on diminishing returns and the junk drawer effect</li>
  <li><a href="https://glenrhodes.com/context-window-management-treating-llm-context-as-working-memory-not-unlimited-storage/">Context as working memory, not storage</a> – Glen Rhodes on treating context like constrained RAM</li>
</ul>

<h3 id="related-posts">Related posts</h3>

<ul>
  <li><a href="/AI/LLM/development/productivity/2025/11/21/orchestrator-automating-claude-code-workflows.html">The Orchestrator: Automating Full Claude Code Workflows</a> – Each phase in its own Claude Code instance</li>
  <li><a href="/ai/development/automation/2026/04/17/automated-builds-cost-fatigue-ceiling.html">Fully Automated LLM Builds: Where It Actually Stops</a> – Token costs as a ceiling on automation</li>
  <li><a href="/ai/development/tools/2026/04/09/gtk-cutting-llm-token-costs-cli-output.html">gtk: Filtering CLI Noise to Save Tokens</a> – Reducing what goes into context in the first place</li>
</ul>]]></content><author><name>Albert Sikkema</name></author><category term="ai" /><category term="development" /><category term="tools" /><summary type="html"><![CDATA[The 1M context window in Claude Code sounds like an upgrade. In practice, constraining it to 200k with early compaction produces better results and lower costs.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://www.albertsikkema.com/assets/images/smaller-context-window-better-claude-code-blog.png" /><media:content medium="image" url="https://www.albertsikkema.com/assets/images/smaller-context-window-better-claude-code-blog.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">What MCP’s Future Means for API Design</title><link href="https://www.albertsikkema.com/ai/development/mcp/2026/04/20/from-talk-to-practice-mcp-future-api-design.html" rel="alternate" type="text/html" title="What MCP’s Future Means for API Design" /><published>2026-04-20T00:00:00+00:00</published><updated>2026-04-20T00:00:00+00:00</updated><id>https://www.albertsikkema.com/ai/development/mcp/2026/04/20/from-talk-to-practice-mcp-future-api-design</id><content type="html" xml:base="https://www.albertsikkema.com/ai/development/mcp/2026/04/20/from-talk-to-practice-mcp-future-api-design.html"><![CDATA[<figure>
  <img src="/assets/images/mcp-future-api-design.jpg" alt="Misty Scottish highland landscape with winding path through moorland" />
  <figcaption>Photo by <a href="https://unsplash.com/@martinbennie">Martin Bennie</a> on <a href="https://unsplash.com">Unsplash</a></figcaption>
</figure>

<p>This weekend I built a small CLI tool that pulls transcripts, comments, and metadata from YouTube videos. The first thing I fed it was David Soria Parra’s keynote <a href="https://www.youtube.com/watch?v=v3Fr2JR47KA">“The Future of MCP”</a> at the AI Engineer conference. David wrote the original Python MCP SDK at Anthropic, so he knows where the protocol is heading. I dove into MCP about 9 months ago as part of the exploratory phase of a government job, a lot has happened since, and there is a lot on the roadmap. I discussed it with an LLM, asking about REST’s role, about playbooks that instruct models how to compose tools, and about the similarities with what I have been building.</p>

<h2 id="what-david-laid-out">What David Laid Out</h2>

<p>The short version: 2025 was about coding agents (local, sandboxed, verifiable). 2026 is about general knowledge workers who need connectivity to five SaaS apps and a shared drive, not a local compiler. He sees three layers for this: skills (domain knowledge in files), <a href="https://modelcontextprotocol.io/">MCP</a> (rich semantics, auth, governance, long-running tasks), and CLI/computer use (great when the tool is already in pre-training data like git or gh). The best agents will use all three.</p>

<p>Three things he wants the ecosystem to fix: <strong>progressive discovery</strong> (stop dumping all tools into context, load them on demand), <strong>programmatic tool calling</strong> (give the model an execution environment to compose multiple calls in one script instead of round-tripping one by one, like <a href="https://blog.cloudflare.com/code-mode-mcp/">Cloudflare’s Code Mode</a> does), and <strong>designing for agents, not REST</strong> (stop mapping REST endpoints 1:1 into MCP servers, he called conversion tools “cringe”).</p>

<p>The upcoming features add a lot, but the new feature I care about most is skills over MCP: servers shipping domain knowledge alongside their tools. More on that below. The protocol is barely 18 months old, with 110 million monthly downloads (roughly 2x faster than React hit that number), so that is proving how popular it is.</p>

<h2 id="i-already-built-this">I Already Built This</h2>

<p>The part that clicked with me was “skills over MCP.” David described it as an upcoming protocol primitive: servers should ship playbooks alongside their tools, instructing the model how to combine them for specific tasks. The server author maintains the playbooks, not the user. When workflows change, the server updates its skills and every connected agent gets the new instructions automatically.</p>

<p>I have been doing exactly this in <a href="/ai/development/operations/2026/04/16/when-llms-actually-deliver.html">logbench</a>, the MCP server I built for querying our Axiom logs. Tools like <code class="language-plaintext highlighter-rouge">explore_dataset</code> don’t return raw data. They return step-by-step instructions: “first get the schema, then run an error breakdown, then drill into the top categories.” The model picks the right workflow tool, gets the recipe, follows it. (Not entirely my idea, did something similar a long time ago, but this step was inspired by Axiom’s official mcp code)</p>

<p>It works well. No context bloat because the playbook only loads when the model calls that specific tool. Progressive discovery is built in for free. And the instructions are scoped to the task at hand, not a generic “here are all the things you could do.”</p>

<h2 id="skills-over-mcp-the-distribution-angle">Skills Over MCP: The Distribution Angle</h2>

<p>Before I get to why formalization worries me, there is one part of skills over MCP that is genuinely exciting: distribution.</p>

<p>Right now, if I want my colleagues to use the logbench playbooks, they need my exact MCP server setup. If I want to share a highly specialized tool behind an auth wall or a paywall, there is no standard way to do that. Skills over MCP solves this. Your team connects to the same MCP server and everyone gets the same playbooks, updated by the server author, no local configuration needed. A specialized log analysis skill, a compliance checking workflow, a financial reporting recipe: all distributed through the same protocol, access controlled at the server level.</p>

<p>That is a real improvement over “copy this markdown file into your project.” It means you can build tools that are genuinely sharable across teams, organizations, even commercially. The distribution story is strong.</p>

<h2 id="the-boring-toolset-problem">The Boring Toolset Problem</h2>

<p>But here is where I am less enthusiastic. What I see happening with MCP is the same thing that happens to every successful protocol: it moves from “wow, that is cool” to the inevitable boring enterprise toolset, the same kind of <a href="/ai/development/2026/02/13/let-the-ai-pick-react.html">standardization convergence</a> I wrote about with React. Mediocre-good-for-all, mostly optimized for large organizations with compliance requirements, not for developers who want to push boundaries.</p>

<p>My concern is not that the formalization itself will limit what I can do. It probably will not. My concern is what happens to developers along the way. When I built the playbook pattern in logbench, I understood exactly what was happening: a tool returns instructions, the model follows them. I learned how to engage with the model, how to structure instructions it would follow reliably, what worked and what did not. That understanding came from building it myself, from prodding and experimenting. Once that becomes a protocol primitive you just consume, the experimentation stops. You get a standard way to do it, and most developers will never look underneath.</p>

<p>That is how we lose the skill of working with LLMs directly. Not because the abstractions are bad, but because they are comfortable. People stop experimenting with how to instruct models, how to structure tool interactions, how to design playbooks that actually work. They use the MCP skills primitive because it is there, and they never discover new patterns that only emerge when you build from scratch.</p>

<p>MCP itself is open source now (Anthropic <a href="https://www.anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation">donated it to the Linux Foundation</a> as part of the Agentic AI Foundation), which is good. But the ecosystem it lives inside is moving in a direction I like less. Claude Code started as a developer-focused tool, the kind of thing where you could <a href="/AI/development/productivity/python/2026/01/13/rethinking-claude-flow-from-per-repo-chaos-to-global-app.html">wire up your own workflows</a> and push the boundaries. Increasingly it is becoming a fits-all product, and the pricing reflects that. The whole Claude Code environment is powerful, but it is also an ecosystem that wants you to stay inside it.</p>

<p>Which is why I think it is worth looking at what exists outside. <a href="https://www.pi.dev/">Pi</a> is one framework worth trying for agentic use, approaching connectivity differently from MCP. There are more options than the one path Anthropic is paving, and the best time to explore them is now, while the patterns are still forming and nothing is locked in.</p>

<h2 id="what-happens-to-rest">What Happens to REST?</h2>

<p>This is the question I kept coming back to. As more and more interaction moves through agents, and agents interact through MCP, what happens to the classic API?</p>

<p>REST is not going anywhere as plumbing. You cannot have MCP without REST (or something like it) underneath. The comments on the talk pushed back hard on this point, and they are right: MCP is essentially “discoverable REST,” and the move toward stateless transport is literally re-converging toward REST patterns.</p>

<p>But I wonder about the trajectory: right now we have API-centered systems and we are moving toward API + MCP. Will that become MCP-centered? Eventually MCP-only for some use cases? And if so, what happens to the APIs that remain?</p>

<p>I think they change character, the classic REST API is a developer’s tool: full CRUD, every resource exposed, every operation available, everything must be in there. The MCP model is different: only those operations that add value, combined into higher-level actions when that makes sense. Less “here are all the building blocks” and more “here is what you can actually do.”</p>

<p>That is not a developer-centered design: it is a human-centered design, or an LLM-centered design, which turn out to be surprisingly similar. And if agents become the primary consumers of APIs, the APIs that stick around will probably start looking more like MCP tools than like the CRUD interfaces we build today. Fewer granular endpoints, more intent-oriented operations. Domain knowledge shipped alongside the API, not buried in documentation.</p>

<p>Or put differently: if your API is so granular that you need a playbook to use it, perhaps the API itself should be the playbook.</p>

<h2 id="what-i-took-away">What I Took Away</h2>

<p>Watch the <a href="https://www.youtube.com/watch?v=v3Fr2JR47KA">full talk</a> if you work with MCP or build agent tooling. It is 18 minutes and dense with where the protocol is heading.</p>

<p>My takeaways, for what they are worth:</p>

<ul>
  <li>The playbook-as-a-tool pattern works today, no protocol extension needed. If you are building MCP servers, try it before waiting for the formalized version.</li>
  <li>Progressive discovery is not optional at scale. If you dump 50 tools into the context window, you are doing it wrong.</li>
  <li>MCP is a good protocol, but keep building things yourself too. The understanding you get from direct experimentation with LLMs is worth more than any abstraction.</li>
  <li>Look beyond Anthropic’s ecosystem. Try <a href="https://www.pi.dev/">Pi</a>, try building without MCP, see what works. The best patterns come from exploration, not from consuming frameworks.</li>
  <li>As agents become primary data consumers, the APIs themselves will loose importance and start looking more like MCP tools: fewer CRUD endpoints, more intent-oriented operations.</li>
</ul>

<hr />

<p><em>Building MCP servers or thinking about API design for agents? <a href="#" onclick="task1(); return false;">Get in touch</a> to compare notes.</em></p>

<h2 id="resources">Resources</h2>

<ul>
  <li><a href="https://www.youtube.com/watch?v=v3Fr2JR47KA">The Future of MCP – David Soria Parra keynote</a> at AI Engineer conference</li>
  <li><a href="https://modelcontextprotocol.io/">Model Context Protocol specification</a> – the official MCP spec and documentation</li>
  <li><a href="https://github.com/jlowin/fastmcp">FastMCP</a> – the Python SDK David called “way better” than the official one</li>
  <li><a href="https://github.com/cloudflare/mcp">Cloudflare MCP server</a> and their <a href="https://blog.cloudflare.com/code-mode-mcp/">Code Mode blog post</a> – example of exposing an execution environment instead of individual tools</li>
  <li><a href="https://www.anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation">Agentic AI Foundation announcement</a> – Anthropic donating MCP to the Linux Foundation</li>
  <li><a href="https://github.com/dsp">David Soria Parra on GitHub</a></li>
  <li><a href="/ai/development/operations/2026/04/16/when-llms-actually-deliver.html">When LLMs Actually Deliver</a> – my earlier post on logbench and the playbook pattern</li>
</ul>]]></content><author><name>Albert Sikkema</name></author><category term="ai" /><category term="development" /><category term="mcp" /><summary type="html"><![CDATA[Watching a talk on MCP's future, I realized I already built the pattern they are formalizing. And it raises a bigger question about how we design APIs.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://www.albertsikkema.com/assets/images/from-talk-to-practice-mcp-future-api-design-blog.png" /><media:content medium="image" url="https://www.albertsikkema.com/assets/images/from-talk-to-practice-mcp-future-api-design-blog.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry></feed>