ETH Zurich tested 100 developers in a controlled, commercial-grade vibe coding environment to see who actually succeeds. The findings cut against both hype camps.
— @thesupermanmx September 5, 2026
Viral framing warning: the study circulated with “China published” and “developers” framing — the authors are at ETH Zurich (Switzerland), and participants were students, not professional developers. The results below come from the paper itself.
The Study
Design details matter because most vibe-coding claims are anecdotes. This one was preregistered (aspredicted.org), run in ETH's Decision Science Laboratory with 100 tertiary students (all with intro CS plus prior LLM-programming experience, C1 English), each in a single 105-minute controlled session for a flat payment. Participants built three GUI apps — a meal-planner replication, a course-scheduler feature addition, and a deliberately opaque “decontextualized” toy app that removes the advantage of recognizing the task — through a custom platform that never showed source code: chat on the left, live preview on the right, with the underlying model being Claude Sonnet 4. Tasks were designed by an eight-expert consensus process across four countries. Scoring was blind human grading on predefined per-feature rubrics.
The Findings
| Relationship | Result | Detail |
|---|---|---|
| CS knowledge → vibe coding | r = .39 (p < .001) | Partial (controlling cognition): r = .281, p = .005 — still significant. |
| Writing skill → vibe coding | r = .29 (p = .003) | Partial: r = .186, p = .066 — not significant after controlling for cognition. |
| Unique variance (hierarchical OLS) | CS ≈ 2× writing | R² .150→.208 adding writing; .083→.208 adding CS. |
| LLM-use frequency → vibe coding | r = −.258 (p = .010) | Exploratory; also r = −.282 vs writing. Causality unknown. |
Reading the table honestly: CS knowledge survives statistical control for general cognitive ability; writing skill does not. In a joint model, CS contributes roughly twice the unique variance of writing — but both add independent predictive value (β_writing = .244, p = .009; β_CS = .356, p < .001). The effect held across the harder task types: writing correlated with the feature-addition (r = .245, p = .014) and decontextualized (r = .239, p = .017) apps, but not the simple replication (r = .154, ns).
Why Writing Matters (and Where It Stops)
The team hypothesized writing would matter most when a task could not be recognized from priors — the decontextualized app. That hypothesis was rejected: the decontextualized correlation was not larger. What did hold is a mediation path through prompt quality (exploratory, cross-sectional): essay skill correlates with human-graded prompt quality (r = .353), prompt quality correlates strongly with vibe-coding success (r = .479), and the indirect path accounts for roughly half the total effect (indirect effect .152, 95% CI [.061, .279]). The authors frame this as response-process validity — clear, structured prompting is writing-as-coding — not as a causal mechanism.
The Frequent-LLM-User Twist
The most-discussed finding: self-reported LLM-use frequency correlated negatively with vibe-coding performance (r = −.258, p = .010) and with writing skill (r = −.282, p = .005), but not with CS knowledge or cognition. Both the paper and ETH's news office stress the ambiguity: heavy LLM use may weaken expression, or weaker writers may simply self-select into heavy LLM use — or both. There was no priming artifact (battery-first vs vibe-first groups were statistically identical). Treat it as a correlation worth taking seriously, not a demonstrated harm.
What This Study Cannot Tell You
Be precise about scope. It is correlational — no causal test that teaching CS or writing improves vibe coding. Participants were students, not professionals, in one city, in one timed session (several ran out of time near completion). The platform showed no code at all, so the CS estimate is a lower bound for mixed code-visible workflows, which may interact rather than scale linearly. The writing instrument was novel (never deployed before), the essay-grading reliability narrowly missed its preregistered target, and the ICAR cognitive scale had low internal consistency. None of this invalidates the findings; all of it bounds them.
Practical Takeaways
For hiring and self-assessment, three evidence-backed signals: (1) CS fundamentals remain the strongest predictor even in no-code workflows — the “prompting replaces fundamentals” claim is empirically backwards at this sample; (2) structured writing predicts task success through prompt quality — invest in specification clarity; (3) familiarity with LLMs is not a skill proxy — frequency of use predicted nothing except, negatively, performance and writing. For tool builders, the authors point to curriculum and tool design: when to emphasize prompt-writing versus CS fundamentals.
Applying the Numbers
Three evidence-backed moves, mapped to the correlations:
- Hire and self-assess on CS fundamentals first (r = .39, survives controls) — a 12-item pseudocode screen predicted no-code success better than any self-reported AI familiarity in this sample.
- Train structured prompting explicitly — the mediation path runs through prompt quality (r = .479 with success), and lexical diversity measures (MTLD r = .343) are cheap proxies you can instrument in your own logs.
- Do not screen for “LLM experience.” Usage frequency predicted nothing — and correlated negatively with both performance and writing. Familiarity is not skill.
For tool builders, the authors' pointer is curriculum design: the CS effect persisted in pure no-code GUI work, so tools that hide code entirely still select for computational thinking — design onboarding and task scaffolds accordingly.
FAQ
Does this prove vibe coding does not work?
No. It identifies who succeeds at GUI-app vibe coding in controlled conditions — and all participants completed most task features. It says nothing about professional, code-visible, or long-horizon workflows.
Should I stop using LLMs to get better at writing?
The study cannot say that — the negative correlation has at least two competing explanations (skill erosion vs self-selection), and the authors explicitly could not determine why. Treat the headline as a hypothesis, not a finding about you.
Which model did participants use?
Claude Sonnet 4, with the study platform built to mirror commercial vibe-coding tools: chat + live preview + rollback, source code never visible.
Is this peer reviewed?
Yes — published at CHI 2026 (ACM Conference on Human Factors in Computing Systems, Barcelona, April 13–17, 2026), with preregistration and materials on figshare.
Sources
- arXiv — Computer Science Achievement and Writing Skills Predict Vibe Coding Proficiency (full paper)
- ETH Zurich News — What skills do people need to successfully program with AI? (official explainer)
- CHI 2026 program entry (publication record, DOI 10.1145/3772318.3791666)
Get the latest on AI, LLMs & developer tools
New MCP servers, model updates, and guides like this one — delivered weekly.