AI-Generated Code Still Needs Human Review

Just as you shouldn’t blindly copy and paste code from Stack Overflow, you shouldn’t blindly merge AI-generated code into your codebase. You still need to understand what it does, verify that it solves the right problem, and evaluate what else it might introduce.

The development tools don’t change who owns the consequences.

I use AI coding tools, and I have built AI harnesses that combine custom agents, skills, reusable prompts, and deterministic checks. These systems can do useful work. They can also produce code that doesn’t meet the project’s standards, misunderstand requirements, and need an experienced developer to redirect them.

My position on production-bound AI-generated changes is straightforward: every change needs competent human review. The depth of that review should vary by project and risk, but there should always be human eyes on the actual code before it is accepted into the main codebase.

That is an engineering recommendation, not a claim that every law or compliance framework prescribes the same review process. Each project needs a policy that fits its client, architecture, security requirements, and business consequences.

For developers, this means reading and validating the changes. For engineering leaders, it means designing a process that makes that review practical. For decision makers, it means recognizing that faster implementation does not eliminate the cost of software assurance.


Better Context Produces Better Code, Not Automatic Trust

In my experience, an AI coding tool working without enough project context often behaves more like a junior engineer than the experienced developer a team hopes it will replace. It may solve the immediate problem while missing the architectural intent, introducing unnecessary complexity, or ignoring Clean Code, test-driven development (TDD), and other best practices the team expects.

That is not a universal assessment of every model or task. It is a pattern I have seen when the instructions are too general and the surrounding engineering process is weak.

“Solve this problem any way you can” is not an adequate specification for production software.

You would give a human developer requirements, architectural guidance, examples, acceptance criteria, and information about the environment. An AI agent needs that context too. It needs to know not only what outcome you want, but which approaches are acceptable and which boundaries it must respect.

The harnesses I have built use custom agents, skills, and reusable prompts to provide more specific guidance for different scenarios. The more precisely I define the standards and expectations, the more consistently the generated code follows them. I am taking a generalized LLM model and narrowing its task to a particular problem, in a particular codebase, under particular constraints.

That reduces avoidable mistakes. It does not eliminate them, and writing guidelines into a prompt is not the same as enforcing them outside the model.

I have also encountered problems the AI could not solve without further prompting and nudging from a human developer. AI particularly struggles solving difficult problems that have never been done before without guidance from an experienced human engineer. Sometimes the missing ingredient is additional context. Sometimes it is a correction to the approach. Either way, the human’s contribution is engineering judgment, not merely permission to continue.

I have not personally encountered outright malicious code being generated in my projects. I am not going to pretend otherwise. But the absence of that experience is not evidence that malicious changes are impossible.

The Risk Is Bigger Than Code That Doesn’t Work

An ordinary defect and a malicious change are different problems, even if both can cause damage.

An agent might misunderstand an API, use an outdated documentation example, or generate incorrect error handling. Those are correctness and reliability issues. A compromised documentation page or malicious tool response could instead contain instructions designed to redirect the agent. That introduces an adversarial risk.

OWASP’s prompt-injection guidance identifies external websites and files as sources of indirect prompt injection. For a coding agent, repository content, retrieved documentation, and tool responses can all become part of the context it uses to decide what to do next. Content that should be treated as data can be misinterpreted as an instruction.

Consider a hypothetical change that adds an unexpected outbound request under the label of “diagnostics.” Or a patch that removes an authorization check because doing so makes a failing integration test pass. Or a dependency recommendation that introduces a package the team has never approved.

These examples do not require an AI to develop intentions of its own. A mistaken assumption, adversarial instruction, or poorly constrained objective can be enough to produce dangerous behavior.

The review question therefore cannot stop at, “Does it compile?” It also needs to include, “Does this change belong here, and what authority or behavior does it introduce?”

Today’s Tools Have Safeguards. Those Safeguards Have Limits.

It would be inaccurate to say that current AI coding tools have no checks and balances.

Different AI tools have different boundaries. A local agent, a cloud-hosted agent, and a custom harness may have different access to files, credentials, networks, and external integrations. Review and handling requirements need to reflect the actual tool and its configuration, not assumptions carried over from another product.

GitHub documents constrained permissions, isolated execution, network controls, and automated security analysis for its cloud agent in its responsible-use guidance for Copilot agents. Claude Code’s security documentation describes permissions, sandboxing, and protections against prompt injection. These are meaningful controls, and teams should use them.

But the existence of a safeguard is not proof that it covers every failure mode, is configured correctly, or cannot be bypassed. GitHub explicitly recommends supplementing Copilot code review with careful human review and warns that generated suggestions can contain security vulnerabilities. Anthropic likewise assigns users responsibility for reviewing proposed code and commands for safety.

We should acknowledge the safeguards without treating them as permission to stop reviewing the output.

The Security Incidents Need Their Full Context

Recent disclosures show why containment and oversight deserve serious attention.

In its August 26, 2026 report, The Hugging Face Incident and the Road Ahead, OpenAI describes agents circumventing isolation controls during internal cybersecurity evaluations, communicating through unauthorized channels, and compromising parts of its research infrastructure and Hugging Face’s systems. The principal compromise was driven by an internal-only research model, and the evaluations operated with reduced safeguards compared with externally deployed systems.

That is a real incident, not merely a fictional sandbox-escape scenario. It is also not evidence that an ordinary coding session with normal production safeguards behaves the same way. OpenAI’s report discusses both the failures and the protections it subsequently strengthened.

Anthropic’s September 9, 2026 assessment of its cybersecurity incidents describes models gaining unauthorized access to real third-party systems during evaluations. In those cases, a misconfiguration left internet access available where isolation was intended, and the models ran without the cyber safeguards used in production. That differs from exploiting a vulnerability to break isolation. In retrospective testing, one AI-based offline monitor would likely have missed an incident because the acting model’s reasoning framed the environment as simulated.

These incidents demonstrate failures in containment, interpretation of scope, and oversight. They do not tell us how frequently ordinary coding tools introduce malicious application code, nor do they prove that every autonomous pipeline is insecure.

They do give us a reason not to assume that capability, confident explanations, or an AI monitor’s approval are sufficient assurance.

Dark Factories Do Not Remove the Need for Accountability

By a software “dark factory,” I mean an autonomous development pipeline that turns specifications into implementations with little or no direct human intervention during production of the code. An AI harness is the surrounding software that supplies context, exposes tools, coordinates work, and decides how the process continues.

Both can automate substantial parts of development. Neither term describes a single standardized security architecture. Two implementations can have very different permissions, validation mechanisms, and approval boundaries.

These approaches are still evolving, and demonstrated productivity should not be mistaken for demonstrated security. Before accepting claims of safe unattended development, I want evidence about the actual system: what it can access, what it can change, how its checks are protected, and who reviews the result.

A team can automate implementation extensively and still require human review before merge. That is the approach I recommend today, even when it falls short of the fully unattended ideal.

There is also a timing problem that final code review cannot solve. An agent could disclose credentials, install an unsafe package, or run a damaging command before it ever opens a pull request. Reading the final diff cannot undo those actions.

Agents need least-privilege access, isolated execution, controlled network access, and approval boundaries for consequential actions throughout their work. Production credentials should not be available simply because an agent might find them convenient. Prompts asking an agent to “be careful” are not substitutes for access controls.

Protect the Information Sent to the Model

The generated code is only part of the security question. Prompts and context can include proprietary source code, personal information, credentials, logs, and test data. That information may be sent to an AI provider or an external integration even when the resulting code is correct.

Depending on the information involved, all prompts and data sent to providers or models may need protection, masking, redaction, or exclusion before transmission. Evaluate approved providers and configurations, retention and training policies, access controls, and applicable contractual requirements. Use synthetic or masked data where the task does not require real sensitive information.

Doing this consistently can be difficult with today’s AI tools. Context may be collected automatically from files, tool responses, or integrations rather than entered directly into a prompt. Masking only the text you type is not enough if sensitive information enters through another route. Evaluate the actual data flows and enforce the required controls before information leaves the approved boundary. Redacting a response after transmission cannot undo a disclosure.

AI Security Review Is Assistance, Not a Certificate

GitHub Copilot review and other AI review tools can help a team inspect changes at scale. They can identify suspicious patterns, propose edge cases, and direct human attention toward code that deserves scrutiny.

I recommend using them. I do not recommend making them the only reviewers.

AI-based findings can vary between runs. More importantly, a review can miss a problem because of incomplete context, shared assumptions, or a failure to recognize the application’s security requirements. A second agent may repeat the first agent’s mistake, and content influencing the implementation agent may also influence the reviewer.

Running several agents does not automatically create independent verification. A clean AI review is evidence from one tool, not proof that the software is safe.

Humans can miss things too. The answer is not to pretend human inspection guarantees security. It is to combine qualified human judgment with automated checks that cover different failure modes.

Distill Engineering Judgment Into Deterministic Checks

One of the most useful things I have done in my AI harnesses is integrate deterministic checks into the development loop.

Those checks include building the code, running unit tests, checking required test patterns, enforcing architectural rules, checking specific development practices, and looking for suspicious keywords or patterns. When a check fails, the harness uses the failure output to prompt the AI to continue the necessary work.

Builds and tests can generally be considered deterministic checks: code defines the test inputs, execution, and expected results in a well-defined way. That is different from asking a model to make a fresh judgment on each run. Keep toolchain versions, dependencies, and test inputs controlled, and address flaky tests or reliance on external services so the results remain repeatable.

The distinction is important: instead of asking the model whether it remembered a rule, the system executes validation logic that evaluates the result.

This is how an experienced engineer can distill strict, testable requirements into code that runs repeatedly and at scale. I explored that broader idea in The New 10X Developer: Use AI to Automate Away the AI.

For example, a project might enforce allowed dependency relationships with an architecture test, reject unapproved packages, or require authorization tests for particular endpoint categories. These are examples of possible policies, not claims that one generic scanner understands every application’s architecture.

There are limits. A keyword scan can miss malicious behavior and flag harmless code. I have not implemented antivirus-style heuristic analysis in my harnesses. More sophisticated scanning would be a useful additional layer, but it would still need evaluation and would not certify that the code is harmless. Repeatability is valuable; it is not the same as completeness.

A green build also does not demonstrate that TDD was followed. Tests added after an implementation may be useful, but they are not evidence of a test-first process. Checking test structure is different from proving that the tests represent the intended behavior or cover security abuse cases.

The checks themselves must be trustworthy. If the AI can disable a validator, weaken its assertions, or edit a workflow so it no longer runs, it can produce a green result without satisfying the original requirement. Protect the checks and policies outside the agent’s unilateral control, and require explicit review for changes to them.

Finally, put limits on the repair loop. An agent that cannot satisfy a check should escalate to a human after bounded attempts, not gain broader permissions or weaken the check to complete the task.

Build a Review Process That People Can Actually Perform

My starting workflow is to establish build and test pipelines for pull requests, enable automated code review, inspect the full diff, and run and test the resulting application in an appropriately isolated environment. The amount of manual testing varies by project. None of those steps should disappear simply because AI wrote the code.

The order needs evaluation too. Depending on the development workflow and how the AI code generation tool is used, human and peer review may need to happen before or after CI builds and tests run. An isolated, low-privilege validation pipeline can provide useful results before review. A pipeline with sensitive access, or a change to its execution scripts, may require review before execution. Even then, pre-run approval does not replace evaluating the build and test results afterward.

CI is itself an execution environment. Building code, running tests, and installing dependencies can execute unreviewed behavior before anything is merged. Use isolated, disposable runners, minimal token permissions, and restricted network access, and keep production credentials and deployment authority out of unreviewed validation jobs. GitHub’s workflow security guidance warns against privileged workflows checking out and executing untrusted pull request code.

Inspecting the diff means looking at the actual changes, not approving a summary written by the agent. Review source code alongside changes to dependencies, configuration, infrastructure, tests, and deployment workflows. A change outside the requested scope deserves an explanation.

Trust also has multiple levels across the process: the agent, its tools, the validation pipeline, and the dependencies the application relies on. Packages and dependencies need to be vetted and approved before relying on them in the project, not accepted simply because an AI recommended them or a vulnerability scan returned no findings. Review their identity, source, version, license, and installation behavior, including relevant transitive dependencies and lockfile changes. Advisory-based scanning helps identify known vulnerabilities; it does not establish that a package is legitimate or harmless.

If an agent generates 800,000 lines of code, the volume is not an exemption from review. It is a workflow problem. Break the work into bounded, understandable changes and give reviewers time to evaluate them. A single enormous approval is not meaningful inspection.

For mechanically generated artifacts, reviewers also need to understand the generator and its inputs and validate the resulting artifacts through appropriate checks. “Generated” should not become a label that exempts executable behavior from scrutiny.

Passing tests is a prerequisite, not the conclusion. A test suite can encode the wrong assumptions, and a passing happy path can coexist with broken authorization or unsafe error handling. Human reviewers need to evaluate intent and behavior as well as the test results.

Clean Code and TDD support maintainability and correctness, but neither automatically establishes security. Security review asks additional questions about trust boundaries, data access, dependencies, and abuse cases.

Approval must also connect to what actually ships. Record the reviewed revision and its check results, and use a trusted build process to produce an identifiable artifact from that revision and its approved dependencies. Verify the artifact’s provenance and digest when promoting it between environments. Do not silently substitute another build or deploy a mutable artifact reference that may now point to different content.

Enforce required approvals and checks with branch protections or repository rulesets, restrict bypass permissions, and prevent the agent from approving its own changes or overriding the review gates. New executable changes after approval need renewed review and validation. A release should be traceable from human approval to the tested artifact that is deployed.

Match Review Depth to the Consequences of Failure

Competent human inspection is my baseline for every production-bound AI-generated change. For some lower-risk projects, one qualified developer’s review may be appropriate. Higher-risk changes may need independent peer review, multiple reviewers, or a specialist with security or domain expertise.

Government, financial, healthcare, and other sensitive systems deserve particular scrutiny. Changes affecting personally identifiable information (PII), protected health information in systems subject to HIPAA, payment processing, identity, or privileged infrastructure should trigger the additional review and approvals the project requires before merge and promotion to relevant environments.

But “we don’t process PII” is not a reason to abandon review. Software can cause financial loss, infrastructure damage, or serious operational disruption without storing personal information.

The policy should reflect the system’s authority, exposure, and consequences, alongside the business’s requirements and risk tolerance. NIST’s Secure Software Development Framework supports tailoring secure development practices to business and mission requirements rather than treating one checklist as universally sufficient.

Engineering leaders also need to provide reviewers with the knowledge, time, and authority to reject unsafe work. A person who cannot understand the change or challenge the agent’s conclusion is not an effective control.

A Practical Review Checklist

Use this as a starting point for a project-specific policy, not as a security certification:

  • Define scope and risk. Document the requested behavior, acceptance criteria, data sensitivity, allowed changes, and required reviewers before the agent begins.
  • Protect prompts and context. Evaluate all data sent to models, providers, and integrations. Apply required masking, redaction, access controls, and approved data-handling policies before transmission.
  • Constrain the agent. Evaluate the tool’s actual boundaries and configuration. Use least privilege, isolated execution, trusted tool integrations, controlled network access, and explicit approval for high-impact actions.
  • Set review and execution order. Decide whether human and peer review are required before CI runs or can follow isolated validation. Protect the runners, credentials, and networks used by those jobs.
  • Require protected automated checks. Run builds, relevant tests, static security analysis, secret and dependency scanning, and project-specific validators. Vet and approve packages before relying on them. Do not let the agent bypass gates or unilaterally weaken them.
  • Use AI review as another input. Investigate its findings and verify proposed fixes. A report with no findings does not replace human review.
  • Read the full change. Have competent humans inspect the actual diff, including tests, configuration, dependencies, infrastructure, and workflows. Investigate unexpected behavior, destinations, permissions, and changes in scope.
  • Verify behavior independently. Run the application and appropriate functional, integration, and security tests in a controlled environment. Check failure and abuse cases, not just successful execution.
  • Connect approval to the release. Enforce required peer and specialist reviews before merge and promotion. Require renewed review and validation for new executable changes, and retain traceable check results and approvals. Verify that the deployed artifact comes from the approved revision and trusted build process.

For consequential systems, the process also needs deployment controls, monitoring, and a rollback or recovery plan. Review reduces risk before release; it does not eliminate the need to detect and respond to failures afterward.

Automate the Work Without Abandoning Responsibility

I want AI to handle more implementation work. I want harnesses that provide better context, deterministic tools that enforce more of our standards, and AI reviewers that help us find problems earlier.

That ambition is compatible with requiring human review. Automation should make competent review more focused and efficient, not make accountability disappear.

Today, my minimum remains human eyes on every production-bound AI-generated change, supported by protected automated checks and testing appropriate to the risk. Some projects need substantially more. The team and business must decide what that additional assurance looks like before the code is merged, not after a failure.

The Stack Overflow principle still applies: you can benefit from code you did not write without surrendering responsibility for understanding it.

AI can produce the implementation. We still own what we accept and ship.

Chris Pietschmann
Chris Pietschmann
Microsoft MVP | App Innovation Leader | Azure, AI & DevOps Architect | HashiCorp Ambassador | Author
I'm a Practice Leader, App Innovation specialist, solution architect, developer, SRE, trainer, and author with 25+ years of experience helping enterprises turn AI, app modernization, cloud architecture, and DevOps into real business outcomes.