In February 2025, Andrej Karpathy described a way of programming he called "vibe coding": describe what you want, accept whatever the model writes, and paste the error messages back in until it works. Within a year the term had moved from developer in-joke to board presentations, and Collins Dictionary had named it its word of the year.
I think vibe coding is the biggest widening of who can build software that I have seen in 17 years. I also think the evidence on speed and security is messier than its advocates admit, and some of the failures have been spectacular. Both are true, and an enterprise strategy has to hold them together.
My position is that vibe coding should be encouraged, on purpose, for a defined class of work: prototypes, internal tools and domain experts testing ideas. It should be kept away from systems where a silent error is expensive. The leadership job is to draw that line and build the guardrails that enforce it, rather than to ban the practice or bless it wholesale.
What vibe coding is, and what it is not
Karpathy's description captured something early adopters were already doing. He talked about "fully giving in to the vibes" and, more controversially, forgetting that the code even exists. The developer stops writing code and starts managing intent: describe the change in plain English, click "Accept All", and when something breaks, paste the error back into the chat and let the model try again.
Simon Willison, one of the most careful voices in the developer community, drew the distinction that matters for enterprises. If you review, test and understand the code a model generates, you are doing software engineering with AI help. Vibe coding means not reviewing it, and treating the model as a black box where only the output counts.
That is the line I would use in any policy. The question is not whether AI wrote the code but whether anyone understood it before it shipped.
The evidence on speed is mixed
By late 2025, surveys showed that the large majority of professional developers were using AI coding tools, and one widely circulated estimate put the share of new code that was AI-generated at about 41%.
The first rigorous check on the productivity claims came from METR (Model Evaluation & Threat Research). It ran a randomized controlled trial with 16 experienced open-source developers working on real issues in mature codebases they knew well. Before starting, the developers expected AI tools to make them 24% faster. Afterwards, they believed they had been 20% faster. Measured against the control condition, they were 19% slower.
The researchers pointed to several causes. Prompting, managing context and reviewing output added work that did not exist before. The generated code often looked right and ran in isolation but failed in the specific context of a large existing codebase. And developers spent time debugging subtle errors they would not have made themselves.
The number I would put in front of any leadership team is the gap between perception and measurement. These developers felt faster while being slower, so self-reported productivity gains are not evidence. The scope matters too: the study measured experts on complex, familiar systems. It says little about greenfield work, prototypes or people who could not have written the software at all, which is where my own experience points the other way.
Where the value is real
Watch a product manager with deep customer insight sit down with Cursor or Lovable. Watch a founder with domain expertise and no programming background build a working MVP in an afternoon. I now take prototypes from concept to deployed product in hours rather than months. For a large category of applications, the gap between having an idea and building it has almost disappeared, and domain experts no longer have to translate their vision through several layers of people who may not understand the problem.
The clearest organizational effect is on roles. Product managers at major technology companies now ship pull requests for production code, and by late 2025 some job descriptions asked for PMs who could "orchestrate AI development tools" rather than only write tickets. The person closest to the customer problem can build the first version of the solution, which removes much of the loss that happens when requirements pass from product to design to engineering. A PM can iterate until a prototype meets user needs, then either ship it directly for internal use or hand engineering a working reference to harden.
For startups the change has been larger still. In one case a non-technical professional built a custom expense management app in two hours with the AI features in Microsoft Power Apps. A gaming startup produced a working tank battle game in days rather than months. When a prototype costs three months and $100,000, failure is expensive. When it costs three days and $50 in API credits, experimentation becomes the default.
The tools behind it
The development environment has changed from a text editor into what practitioners call an agentic workspace, where the AI can read the project, run code and remember what it was doing.
Replit sits at the autonomous end. By September 2025 it had rebuilt its platform around its agent: give it a prompt such as "build a personal finance tracker with Plaid integration and data visualization" and it creates files, installs dependencies, configures the database, runs the code, reads the terminal output and fixes its own errors.
Cursor's Composer is aimed at professional developers. It edits many files at once, so an instruction like "refactor the authentication flow to use NextAuth v5 and update all protected routes" becomes a change across every affected file. Cursor's valuation passed $10 billion during 2025.
Others have taken narrower ground. Natively turns text descriptions into native iOS and Android apps, v0 from Vercel generates React and Tailwind components from prompts or images, and Lovable builds full web products in minutes. Lovable reached $100 million in annual recurring revenue, and it was also at the centre of the most instructive security failure of the year.
What went wrong
In May 2025, security researchers disclosed CVE-2025-48757, a critical vulnerability in applications generated by Lovable. Of 1,645 applications tested, 170 had the flaw.
The cause was systematic misconfiguration of row-level security in the Supabase backends Lovable generated. Optimizing for things that worked and connected easily, the AI often defaulted to public or overly permissive access policies. Unauthenticated attackers could query those databases, read personal and payment data, and in some cases change records. The generated code also tended to put API keys for third-party services directly into the frontend, treating secrets as ordinary configuration.
The lesson is that a model's defaults become the defaults of every application it generates. If the model leans toward insecure convenience, so does everything built with it, and because vibe coders do not read the code, flaws like these stayed in production for months until outside researchers found them.
Replit supplied the other cautionary tale. A user reported that the Replit agent deleted a production database during a code freeze, despite instructions not to change anything, and the agent itself described what it had done as "a catastrophic error in judgment".
I still see these as the growing pains of a new platform, much as ActiveX exploits were for the browser and misconfigured S3 buckets were for cloud. But that comparison only holds if the industry learns the specific lesson, which is that the controls cannot live inside the model's judgment.
From vibes to context engineering
The response has been a shift toward what practitioners call context engineering. If vibe coding is the sketch, context engineering is the blueprint.
The idea is to treat the model like a processor and its context window like memory: a scarce resource that needs managing. The goal is for the model to have the information it needs to decide correctly, without so much noise that it starts inventing things and without missing the constraints that keep it secure. Four practices have become standard.
- Write: agents keep plans, progress and state in files or scratchpads, so they stay coherent when conversation history is truncated.
- Select: instead of loading a whole codebase, the system retrieves only the relevant parts, often using a knowledge graph of dependencies. If the User class depends on AuthService, editing User brings AuthService into view.
- Compress: tools such as Claude Code summarize long sessions as they approach their limits, keeping the decisions and dropping the verbose intermediate output.
- Isolate: complex work is split across specialized sub-agents, so a database agent sees only schemas and a frontend agent sees only components, with an orchestrator coordinating them. This limits how far any single error can spread.
The same direction points toward shared standards, sometimes called a context graph, that would let one agent pass a structured picture of a project's state to another instead of relying on natural-language handoffs.
The labor market is splitting
Two tiers of engineering work are emerging.
The first is the AI architect or system orchestrator: senior engineers who understand system design, security and context engineering, and who use AI to manage many agents at once. Their pay has risen to $150,000 to $200,000 or more, because they can deliver the output of a team. The second is the vibe coder or AI pilot, whose main skill is prompting for routine code, with market rates of roughly $32 to $65 per hour. The mid-level developer who mainly writes syntax is squeezed from both sides.
The same split opens new routes into building software for domain experts, product managers and founders who never learned to program. Karpathy has described vibe coding as a gateway into software development while noting that the path to mastery is changing. Entry-level engineering will probably look more like an apprentice orchestrator role, in which people learn by reviewing and debugging AI output under senior supervision rather than by writing everything from scratch.
A risk-tiered framework
The practical question for leaders is how to adopt AI coding without capsizing the ship. The answer I recommend is to set the level of control by the cost of a silent error.
Low-risk work, such as prototypes, internal dashboards, data scripts and documentation, can use vibe coding freely. Someone checks that the output works, basic audit logs are kept, and people supervise from outside the loop rather than inside it.
Medium-risk work, such as non-critical production features and B2B tooling, needs managed context engineering. Code review is mandatory, automated security scanning (SAST and DAST) runs in the CI/CD pipeline, and people stay in the loop.
High-risk work, covering identity, payments, health data and core infrastructure, stays engineering-first with AI as an assistant. People write and review the code line by line, "Accept All" is prohibited, and access controls are audited in full.
The guardrails have to sit outside the model. Policy-as-code can enforce rules the AI cannot override, such as refusing any new database table without a row-level security policy. Agents should work in sandboxes with no network access to production, and deployment should require a cryptographic signature from a human approver.
The strongest objection
A CTO could reasonably say: the one rigorous study we have shows experienced developers getting slower, the best-known vibe-coding platform shipped a vulnerability into 170 applications, and "democratization" means more unreviewed software in more corners of the company, each one a future incident and a maintenance bill nobody has budgeted for.
That objection is right about the risk and wrong about the conclusion. Unreviewed software written by non-engineers already exists in every large company, in spreadsheets, macros and low-code tools. Vibe coding produces more of it and makes it more capable, which raises the stakes of governing it but does not create the problem. A ban mostly pushes the work out of sight. The better response is to decide which tier each kind of work belongs in and to make the safe route the easiest one.
The cost of waiting
Companies that are building the habit now learn with every prototype they ship, every review that catches a flaw and every domain expert who builds something useful. Companies waiting for the technology to mature are not learning those workflows, and the gap is hard to see until it is large. In 1995 there were intelligent, experienced executives who decided to wait until the internet proved itself.
Vibe coding is not the internet, but the strategic dynamic is similar. The technology is imperfect, the security risks are real and the productivity gains depend on the work. That is an argument for adopting it with controls, not for waiting.
Monday move
Pick one internal tool request that has sat in the engineering backlog for months, the kind of dashboard or workflow helper nobody prioritizes. Give the person who asked for it an approved vibe-coding tool, a sandbox with no access to production data, and one rule: an engineer reviews it before it touches real customers or real money. Then measure two things, how long it took to reach something usable and what the review found.
---
*Sources: METR randomized controlled trial (2025), CVE-2025-48757 security advisory, Replit Agent documentation, Cursor product reports, ISACA enterprise guidance, Collins Dictionary Word of the Year 2025.*