Getty Images/iStockphoto

Tip

The risks of AI overreliance in software engineering

It's unquestionable whether generative AI provides productivity gains for software engineers. But developer overreliance on AI comes with its share of risks.

AI coding assistants have advanced beyond autocomplete. Current coding agents can inspect repositories, change multiple files, execute commands, run tests and continue working after encountering errors. The amount of work that engineers can delegate to coding assistants is expanding quickly.

In the paper "Measuring AI Ability to Complete Long Software Tasks," AI model researcher METR found that the frontier systems' task horizon -- how long an autonomous agent can operate without human intervention -- has historically doubled about every seven months since 2025. The benchmark tasks METR assigned don't encompass the full complexity of production engineering, but the trend is clear. Software tasks that were beyond agents' capabilities only a few years ago are increasingly within their reach, and METR researchers believe that within a decade, agents could be completing "a large fraction of software tasks that currently take humans days or weeks."

Large input context windows in models from developers like OpenAI, Claude and Gemini, along with longer output limits, have enabled AI assistants to ingest large quantities of source code, documentation, test results and conversation history during a coding task. These capabilities are some of the drivers that make these models more efficient at completing engineering tasks.

However, this efficiency could easily instill overreliance on these tools in our developers, leading to time-consuming issues; problems with code ownership and explainability; and, at worst, errors in business-critical software.  

Faster code changes the mix of engineering work

Engineers increasingly supervise code they didn't personally produce.

A 2026 Management Science study, "The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers," examined software developer productivity at Microsoft, Accenture and an anonymous Fortune 100 company. Of the 4,867 developers surveyed, the study found that those who had access to an AI-based coding assistant completed 26% more tasks. Less experienced developers showed particularly high adoption and productivity gains.

A 2026 paper, "A meta-analysis of the effect of generative AI on productivity and learning in programming," found that across 23 programming studies, researchers reached a broadly similar conclusion to the Management Science study. GenAI produced a moderate positive effect on programming productivity, although results varied substantially by setting. However, there was no evidence that the use of GenAI improved learning outcomes or skill development.

The lesson for CIOs, CTOs and IT leaders is that AI provides engineers with more implementation capacity, and this increased capacity changes the economics of engineering while also shifting bottlenecks. Requirements, architecture, testing, review, deployment and operation don't necessarily accelerate at the same rate. And as AI development itself becomes cheaper, testing, verification and deployment account for a larger share of the remaining work. AI can reduce total engineering effort, but it also changes the composition of that effort.

Graph illustrating how AI is changing the share of work that software engineers are completing.
AI can reduce total engineering effort while changing its composition. As development becomes cheaper, testing, verification and deployment account for a larger share of the work that remains.

A 2025 DORA study, "State of AI-assisted Software Development," describes AI as an amplifier of an organization's existing engineering systems, magnifying the strengths of high-performing organizations and the dysfunctions of others. Its 2026 qualitative work "The ROI of AI-assisted Software Development," found that time saved during creation was frequently reallocated to auditing and verification, what DORA calls the "verification tax." Higher AI adoption was associated with higher delivery throughput and higher delivery instability. That changes the meaning of developer productivity. Producing an implementation faster is valuable, but engineers must still understand, review, test, secure, operate on and update that implementation.

A July 2026 study "AI Writes Faster Than Humans Can Review: A Longitudinal Study of an Enterprise '2 x' Mandate" details what can happen next. After adopting AI coding tools, the code reviewer load doubles, and automated review overtakes human review. The study points to a practical constraint: Software generation can scale faster than human review, shifting the verification bottleneck downstream.

AI-generated code creates an obligation for businesses to certify its correctness. For routine boilerplate work, the effort can be small. Authentication, payment systems, concurrency, infrastructure and database migrations impose a different verification burden.

In Stack Overflow's "2025 Developer Survey" of more than 49,000 respondents, 66% of developers reported frustrations with AI tools that were "almost right, but not quite," while 45% said debugging AI-generated code was more time consuming. Plausible code deserves particular attention. An implementation can use familiar patterns, sensible names and convincing comments while containing a subtle error.

The capability at risk is understanding

A 2026 study in IEEE Transactions on Software Engineering, "More Code, Less Understanding? On the Impact of AI Assistants on Developers' Productivity and Code Ownership," provides evidence of the trade-off between developer efficiency and ownership, or the ability to argue about implementation choices, when using AI tools in engineering. Researchers gave 69 participants code writing tasks with and without AI assistance. Participants using AI achieved more than twice the median task completeness, but their ability to answer technical questions about the code lowered by 12.5%.

To be sure, engineers don't need to retain every implementation detail. Programming has progressed through layers of abstraction. Libraries, frameworks, compilers, IDEs and search engines all reduced the amount of information developers had to remember.

Program comprehension, however, is different. An engineer responsible for a system needs to understand where state lives, how data moves, which assumptions the components make, how dependencies fail and where security boundaries sit. Debugging sharpens those skills. A skilled engineer forms a model of expected behavior, develops hypotheses, inspects evidence, eliminates explanations and validates the eventual repair.

AI can strengthen that process by suggesting hypotheses, interpreting logs or proposing experiments. It can also produce a different workflow: paste an error into the model, apply the proposed fix, paste the next error and apply another fix. The application might work, but without debugging, the developer's understanding will suffer.

Evidence of an industry-wide decline in professional debugging skills doesn't yet exist. Generative coding systems haven't been deployed long enough to establish such a longitudinal effect. But the general mechanism of deskilling is well understood, and research on automation has found that removing routine activity can also remove opportunities to maintain the skills humans need when automation encounters edge cases.

The issue becomes particularly important for new engineers entering the profession. Junior developers traditionally acquire expertise through ordinary work: small bugs, tests, unfamiliar code, failed implementations, documentation and code-review feedback. AI is increasingly capable of handling those same tasks. The first cohort of engineers beginning their careers with modern coding agents has just begun, and research will show what 10 years of AI-assisted career development will produce.

Employers, therefore, need to consider the learning value of work before automating it away. Asking AI to explain why a programming flaw occurs exercises a different capability than asking the system to fix it. Both approaches can be efficient, but their effects on skill development differ.

Measure the engineering system, not the volume of AI use

High AI usage alone provides little evidence of overreliance. An experienced engineer can use an agent throughout the day while retaining a strong understanding of the system. A more useful warning sign appears when dependency grows while independent engineering capability falls.

Executives can see parts of that change in existing operational data. AI adoption and pull request volume can rise even as review queues lengthen. Defects, reverts or incident-recovery times might increase. Junior engineers might produce sophisticated implementations while struggling to explain their behavior.

Code ownership can also be tested directly. An engineer responsible for an important AI-generated change should be able to explain why it exists, how it works, how it can fail and why the evidence supports its correctness.

Matrix showing how AI-generated content scales with human understanding.
AI-generated output is most valuable when human understanding scales with it. High generation combined with weak comprehension creates a control problem even when short-term output looks strong.

AI-generated output is most valuable when human understanding scales with it. High generation combined with weak comprehension creates a control problem even when the short-term output looks strong.

Teams can use agents aggressively for boilerplate, test scaffolding, routine transformations, exploratory work and well-bounded implementation. But important changes still need accountable owners. Critical requirements should be expressed before an AI system generates code and tests. Compiler checks, static analysis, security tooling, human review and production telemetry provide different forms of evidence and should remain distinct.

Junior engineers need opportunities to debug, review and reason without delegating their entire task. Periodic work without generative assistance can reveal whether the organization still possesses the independent capability required to handle difficult failures.

AI has already made software cheaper to generate. The current trajectory of coding agents, large context windows and enterprise deployments suggests that this trend will continue. Human understanding has no comparable scaling mechanism. As the volume of generated software rises, companies will depend on engineers who can determine whether that software is correct, trace what happens when it fails and take control when the agent reaches the edge of its competence.

AI can materially increase engineering output, and coding agents are likely to take on a growing share of implementation work. The risk is that generation capacity advances faster than the human ability to understand, verify and maintain what is produced. Organizations will need to measure engineering outcomes beyond code volume, preserve independent technical capability and ensure that greater AI use does not weaken control over critical business software.

Kashyap Kompella, founder of RPA2AI Research, is an AI industry analyst and advisor to leading companies across the U.S., Europe and the Asia-Pacific region. Kashyap is the co-author of three books, Practical Artificial Intelligence, Artificial Intelligence for Lawyers and AI Governance and Regulation.

Next Steps

Dig Deeper on AI Ethics & Governance