“A lot of them has moved to step 3 and 4. Some of them also to 5 running full agentic mode… I have one developer stating that I haven’t wrote one single line of code the past month. So it’s really getting under the skin on the developers.” A senior manager said that about his own rollout, success and damage in the same breath, describing a 1,200-employee Danish software firm — called SoftHouse in a case study published 3 September 2026 — that built a 1-to-5 AI maturity model, level 5 meaning full agentic coding, and pushed staff to climb it. A year in, the firm’s comprehensive AI policy was still unwritten, and which tools were permitted varied by customer contract. Practitioners were graded on how much AI they used without being told what they were allowed to use it for.
The ladder graded usage the policy hadn’t caught up to
SoftHouse serves regulated clients in health, taxation and public administration, and its maturity model was explicit: practitioners rated their own usage against five levels, “with organizational ambitions for practitioners to reach the upper levels.” P19 put it bluntly: “they defined maturity levels, and up to a level 5… so they’re pushing us to evaluate us.” With the policy on permitted tools still being written, many practitioners froze rather than risk it. P11: “I have heard people say: I don’t know what I’m allowed to do, so I don’t do anything, to not run the risk that I do something I was not allowed to do.”
That’s the mechanism, not just a mood — a pincer, in the researchers’ framing. The firm’s “trust-by-default stance converts every practitioner into their own compliance officer,” pushing liability onto individuals with no guardrails to judge against. If your employer runs a similar programme, this is what’s at stake for you: the level you enter into your own self-assessment is a number your manager reads — and a developer three years in, filling out that form each quarter, is guessing at rules that were never written down. Grading how far someone has climbed a ladder is not the same as telling them which rungs are safe to stand on.
Grading how far someone has climbed a ladder is not the same as telling them which rungs are safe to stand on.
Advanced users describe reinvention, not loss — for a narrower group than it sounds
The counterargument has real weight. GitHub’s own developer research , published by Eirini Kalliamvakou, builds a similar four-stage fluency ladder — Skeptic, Explorer, Collaborator, Strategist — and reaches the opposite emotional register: advanced practitioners describing a shift “from code producers to creative directors of code,” not a loss of craft. It’s a fair reply, and it deserves to be taken seriously.
But look at who GitHub recruited: 22 engineers already using AI for more than half their coding work, with hands-on experience across at least four tools. That filter selects for people who’d already made peace with the tool — it structurally excludes anyone a ladder might be pressuring, since they wouldn’t have qualified for the study. SoftHouse’s own authors are careful about their case’s limits, too: they write that the psychological costs they found “coexist with acceptance” of AI and are “not evidence that they resist adopting it,” and they don’t claim the pattern “will transfer unchanged to other organizations.” It’s one firm, cross-sectional, in regulated sectors where compliance anxiety runs high, with 19 of 21 participants a decade or more into their careers — a real constraint on how far the finding travels, not a reason to discount it.
One company’s ladder, but the shape is now the industry default
None of that scopes the finding down to SoftHouse alone, though — the instrument’s shape has already spread. SEI and Accenture released a five-level organizational AI adoption model in June 2026, built from executive interviews and surveys of nearly 600 practitioners, scoring companies rather than people across dimensions that include Risk and Governance. Neither that model nor GitHub’s fluency stages was the subject of SoftHouse’s study; extending the finding from one Danish firm’s usage ladder to those instruments is a step this article is taking, not one the researchers measured. It’s an imperfect extension too: GitHub’s ladder grades fluency, not the usage level SoftHouse tracked, so the two aren’t measuring identical failure modes even though they share rungs. It’s the trap Amazon and Meta hit with internal AI-usage leaderboards before scrapping them: whatever you grade people on stops measuring anything once the grading becomes the point.
What SoftHouse’s paper does establish, inside its own case, is a specific prescription: “clarify governance boundaries, decouple adoption expectations from individual measurement, and recognize verification work as legitimate effort.” That third clause matters as much as the first two — it echoes a governance case study finding real guardrails discovered from failures mid-project, not specified upfront , exactly the sequence SoftHouse got backward: the ladder arrived before the rules did.
The manager quoted at the top wasn’t wrong that his rollout worked — developers really did climb to level 5. He also wasn’t wrong that it was getting under their skin. A maturity model that can’t tell those two facts apart isn’t measuring adoption. It’s measuring how much pressure a number can generate before anyone asks what the number was supposed to mean.



