Miggi Cardona, a designer advocate at Figma , pitches the company’s new skill-authoring feature by picturing a colleague whose future work comes out with “motion designs eased like mine” — your taste, running in someone else’s hands, according to Figma’s blog . On 13 August 2026, Figma shipped exactly that, confirmed the same day in its own release notes . The pitch is almost entirely about personalization — skills, Figma says, are “perfect for capturing the opinions and judgment calls that every designer carries around.” Three days earlier, the first study to actually test that premise found the opposite: personalized instructions barely moved an AI agent’s performance, while instructions shared across everyone worked measurably better.
What a skill is, and what Figma is actually claiming
A skill, in this new vocabulary, is a plain markdown file — a SKILL.md, in the format Figma and coding tools like Claude now share — giving an agent ordered instructions instead of leaving it to guess. You ask Figma’s agent to distill one from a design frame, test it in chat, then publish it. Figma’s Community library already holds more than 50, covering research, design-system conventions and handoff — the point, as Mazette’s coverage puts it, is making sure “every new prototype follows the same standards for components, naming, and tokens” without re-explaining the rules each time. That library is the generic option. The marketing energy goes toward the bespoke one — the skill encoding your particular judgment calls, not the shared defaults everyone downloads.
The instructions everyone shares beat the ones written specifically for you, and the study's authors can't yet say why.
The study measured coding agents, not design agents
That marketing bet is worth testing against “Do Personalized Skills Help Coding Agents?” , submitted 10 August 2026, which ran 206 real developer sessions with a coding agent (Codex on GPT-5.5) across 13 developers. Skills distilled from an individual’s own interaction history produced a score of 65.99 against a no-skill baseline of 65.02 — a gain of 0.97, not statistically significant. Skills built by pooling everyone’s sessions together scored 68.80, a gain of 3.78 — larger and more consistent than the personalized result, but one the authors say still falls short of the conventional significance threshold (paired t-test, p=.063). In head-to-head matchups the personalized file actually lost to having no file at all more often than it won — 41.43% wins against 43.81% losses — while the pooled generic skill won 50.95% of the time and lost only 34.29%. The bespoke instructions were, more often than not, worse than nothing.
| Category | Score |
|---|---|
| No-skill baseline | 65.02 |
| Personalized skill | 65.99 |
| Pooled generic skill | 68.80 |
That result was measured on coding agents distilling procedural habits from thin session logs, not on a designer authoring a skill in Figma from a frame they already understand. Reading it across to Figma’s launch is a real step, not a given: the study shows automatically inferred personal preference underperforming shared procedure in one narrow setting — code review, feature work, testing and infrastructure tasks, scored by an agent. Applying it to hand-authored design skills means trusting that thin signal losing to broad pattern generalizes past the exact mechanism the paper tested.
The study’s own limits cut against overclaiming, not the study’s finding
The paper’s authors are upfront that 206 sessions from 13 developers “limits our ability to draw definitive conclusions,” and their own data shows personalization becoming effective once a developer had six or more relevant historical sessions — a threshold their dataset rarely reached. That matters doubly here, because the study’s skills were auto-distilled from sparse logs, while Figma’s are deliberately written by someone who already knows what the convention means. The mechanism that lost in the paper isn’t quite the one Figma shipped — and the same practitioner coverage notes the Community library carries its own risk: “a poorly designed skill … can introduce bad habits just as easily as a good skill fixes them.” Personalization and generic skills fail in different ways; neither wins outright yet.
None of that erases what the numbers say today. A design-system lead about to spend a sprint packaging house conventions into SKILL.md files, and a developer maintaining a hand-tuned CLAUDE.md, are both being sold the same premise Figma is selling — that encoding your own judgment beats the shared default. On the only evidence anyone has measured so far, the boring pooled skill should go in first, and the bespoke one has to prove it beats writing nothing at all, which the study says it often doesn’t. This is also where design systems’ broader reliability problem with AI agents reappears: a rule written down is not the same as a rule followed, whether that rule came from one designer’s taste or fifty.
Figma’s own metaphor for a skill is closer to a recipe than a rulebook — something handed off and expected to work the same way twice. The evidence so far says the shared community recipe, the one nobody wrote for you specifically, is the one actually working in the kitchen; your own handwritten card, however carefully kept, is still unproven. That is the same tension already visible in forkable brand-identity files other teams have started sharing : the generic version travels well because nobody had to guess what it meant.



