# AI Code Review on GitHub: Automating PR Review Without Writing a Test Suite

Get behavioral feedback on every GitHub pull request before you write a test suite. Learn how to set up runtime PR review and where human judgment still decides the merge.

By Evan Marshall · 2026-10-02

Every senior engineer runs the same mental exercise when a PR lands in their queue. They scan the diff, translate the change into a list of things that could break, and wonder whether the author tested any of them.

Dozens of times a week, reviewers mentally work out what a code change could break and what they need to test. It's expert labor, and it's almost entirely invisible, and until recently the main way to make it reusable was to author a test suite upfront.

But writing, maintaining, and populating an environment for a full suite takes time that many teams don't have before they need feedback.

Reviewing code, proposing a test plan, and executing that plan are three separate activities.

A tool that reads the diff and leaves inline comments handles the first activity. A tool that translates PR context into scenarios handles the second. A tool that builds and runs the application against those scenarios handles the third.

Knowing which of those your team needs, and where the human reviewer still belongs, makes the difference between useful automation and a green badge that doesn't mean anything.

In this guide, we’ll cover a GitHub-specific workflow, a shortlist of 7 compatible tools, and a concrete example that connects a code change to observed behavioral evidence.

By the end, you'll know what to connect, what context to supply, what actually runs, and what stays a human decision to help you approve PRs and merge them with more confidence, faster.

## How to automate GitHub PR review without a test suite

You can automate parts of PR review without writing a test suite first. But checking how a change actually behaves requires a runnable application and a plan for testing it.

Building that plan is the exercise you've been running in your head on every PR, and you can hand it off now. You still make the merge call, but you no longer have to come up with the test scenarios yourself or write a suite before anyone can check the change.

Here's how the workflow runs.

1. Connect your GitHub repository to the review tool.

2. Write PR descriptions that explain intent, expected behavior, and scope, so the tool has material to work from.

3. Configure the environment your application needs to run, including credentials, seed data, and service access.

4. Trigger a review, observe the run state, and read the findings with reproduction evidence.

5. Use the results alongside your existing CI checks and human approval to make the merge decision.

The sections below walk through each step using [Ito](https://ito.ai/docs), which generates a targeted test plan from PR context and executes the changed application in an isolated environment. That walkthrough answers the problem-solving question directly, showing how you can catch runtime bugs before merging a GitHub pull request even without a pre-written test suite.

Review tools that read your code can catch logic errors, naming problems, and pattern violations in the diff. Tools that run your application can catch behavioral regressions, data integrity failures, and broken user flows that only appear when the code executes.

Both signals matter, and the setup steps, latency, and evidence each approach returns differ, which is why the tool list below covers both.

## What are your options for AI code review on GitHub?

Most AI review tools for GitHub read the diff and leave comments on it. Ito, our product, builds the changed application and runs a generated test plan against it.

The tools below connect to GitHub in 3 ways. Copilot is built into GitHub, most others install as hosted GitHub Apps, and ai-review runs as an Actions workflow you maintain yourself.

Before you connect a tool to your GitHub repository, here's what to expect from each option.

|  | Tool | How it connects | What you get back | Worth knowing |
| --- | --- | --- | --- | --- |
|  | Ito | GitHub App | Findings from a running build, with video for web apps and scripts for backends | Needs credentials, seed data and a base URL. Reviews typically take 30 minutes to 2 hours |
|  | GitHub Copilot code review | Built into GitHub | Inline suggestions plus an overview comment | Not on Copilot Free. Uses Actions minutes on private repos |
|  | Nikita-Filonov/ai-review | Self-managed Actions workflow | Inline, summary or cross-file comments from the model you choose | You own API keys, runner costs and fork security |
|  | CodeRabbit | GitHub App | Summaries and inline findings, configured in .coderabbit.yaml | Hourly review limits vary by plan |
|  | Qodo | GitHub App | Severity-ranked findings with fix guidance | 14-day trial. PR-Agent is a separate open-source project |
|  | Cursor Bugbot | Cursor's GitHub integration | Findings with "Fix in Cursor" links | Findings report as neutral by default, so the check alone won't block a merge |
|  | Claude Code Review | Anthropic's GitHub App (managed) | Inline findings tagged Important, Nit or Pre-existing | Team and Enterprise only, research preview, about $15 to $25 per review |

For setup steps, pricing, and a fuller comparison, see our roundup of AI code review tools for GitHub.

## Set up Ito for automated GitHub PR review

The 5 steps below follow one repository from GitHub App connection through test planning, environment configuration, findings review, and the merge decision.

You'll move between GitHub's own installation flow and our dashboard along the way, but there's no review-workflow YAML to author and no test scripts to write before the first run.

### Step 1: Connect your GitHub repository to Ito

Go to [app.ito.ai](http://app.ito.ai) and select "Sign up with GitHub." Ito uses GitHub OAuth for authentication, so no separate password is needed. After authentication, the onboarding wizard opens at /auth/setup.

In the Connect step, you'll see your existing Ito installations if any exist. Select one, or click "Install App" to go to GitHub's installation page. Choose the organization or personal account, then grant access to specific repositories rather than all repositories in the organization. Click Install, then Continue back in the Ito wizard.

Two things can slow this step down.

If you don't have admin or owner rights on the GitHub organization, GitHub queues the install request instead of completing it, and onboarding shows "Waiting on organization owner approval" until an owner approves it. And because onboarding lists only your own open or recent PRs, you'll need at least one in the repository you're connecting to, ready to be built and used for testing.

Once onboarding completes, confirm the repository appears in the Repositories list in our dashboard before moving to settings.

For an existing installation where you need to adjust repository access later, go to Repositories > Manage GitHub Access. An authorized GitHub administrator can add or remove repositories there without repeating the full onboarding flow.

This integration uses the [Ito GitHub App](https://github.com/apps/itoqa), so there's no review-workflow YAML to maintain.

### Step 2: Give Ito context for an AI-generated test plan

Ito builds its test plan from the diff and the PR description, so the plan is only as good as the context you give it. Write the description for someone who has to test the change, not just review it: explain what the change does, how the affected flows should behave afterward, and what's deliberately out of scope.

We don't read issue trackers directly, so ticket context needs to reach the PR description one way or another.

If you use Linear, linking an issue by putting its ID in the branch name or PR title, or writing a magic word like Fixes ENG-123 in the description, is enough. [Linear's GitHub integration](https://linear.app/docs/github-integration) then posts the issue's title and description as a PR comment, which we read as additional context. For other trackers, pasting the relevant ticket details straight into the PR description is the reliable path.

For priorities that apply across every PR, rather than only one, use [Test Instructions](https://www.ito.ai/docs/settings/test-instructions) instead. A line like "always verify authenticated flows" or "skip the admin dashboard for now" applies org-wide by default and can be scoped to a specific repository or PR author.

Treat it as a starting point rather than a finished configuration: Ito calls this a new and evolving capability, and plans improve as your team closes findings as won't-fix and Ito learns from what gets fixed.

To see what good context produces, take [DoltgreSQL PR #2913](https://github.com/dolthub/doltgresql/pull/2913). The PR moved window function support into the Doltgres layer specifically so its behavior would match PostgreSQL's. A description stating that intent, plus the expected SQL output, gives Ito exactly the material it needs to generate targeted scenarios instead of generic ones.

### Step 3: Configure the environment and run your first review

Before Ito can execute anything, it needs to run your application the way a person would. Open [Settings > Context & Secrets](https://app.ito.ai/settings/context-secrets) in our dashboard, select the repository, and supply what it needs:

- Variables for non-sensitive configuration like a base URL or feature flags.

- Secrets for credentials, API keys, and private-package tokens. Use dedicated test accounts and nonproduction credentials: our tests can exercise adversarial behavior, so credentials that touch production data don't belong here.

- Seed data, if the application needs specific records to reach the states the test plan exercises.

With that in place, open [Settings > Automation](https://app.ito.ai/settings/automation) to turn reviews on for the repository.

To change these settings, you need to be a Github organization admin, an Ito Manager, or a personal-account owner. Reviews only run for PRs opened by active team members, so check that the PR author qualifies before you turn Reviews on.

We post a summary on every PR where a run completes, including all-pass runs, with a message like "15 test cases ran, 15 passed." Leave Comments off and results stay in the dashboard only.

From here, opening or updating an eligible PR is all it takes. Then, you can watch its state move from Pending to Running in our dashboard.

Each run is tied to a specific commit SHA, so pushing a newer commit while a run is executing cancels the stale run and starts a fresh one against the new commit.

> **An important note:** a build failure or an inaccessible service isn't a behavioral finding; it's an environment problem. A run that can't build the application or reach a required service returns no evidence either way, so you should treat that outcome as something to fix, not a pass.

returned 30 for id 1, while the equivalent inline query

returned 10 for the same row. Under PostgreSQL semantics those two queries should return identical results, and the inline query's output was the correct one. DoltHub had test coverage for this syntax already, it was just asserting the incorrect result, which is exactly the kind of gap that reading the diff alone won't surface.

Findings like this one, and Additional Findings, both carry one of 4 severities:

- **Critical:** Core functionality or key user flows are blocked, or data integrity is at risk.

- **High:** Important features are significantly degraded, but the application isn't completely blocked.

- **Medium:** Usability is affected, but core functionality remains accessible.

- **Low:** Limited impact, often cosmetic or triggered only in rare conditions.

Severity doesn't decide whether a run passes: a run fails if any PR-attributed test case fails, regardless of that case's severity, so a failed run can still be a Low-severity bug. As a rule of thumb, fix Critical issues before merge and treat High issues as usually worth fixing before shipping as well.

Additional Findings sit outside that pass/fail calculation because they're pre-existing issues, not something the current PR introduced. That doesn't mean ignore them. A Critical pre-existing issue is still worth assigning an owner, even if it isn't blocking this particular merge.

When you push a fix, Ito starts a new run against the new commit. Compare that run's results to the one before it, and use the Origin labels, New, Regression, Still broken (verified), Still broken (inherited), plus the separate Bugs Fixed section, to see at a glance whether the fix actually landed.

Here’s how the diff summary from Ito would look like on a PR after your commit fixes all the bugs:

### Step 5: Use Ito's results in your GitHub merge decision

Ito's review informs the merge decision; it doesn't replace GitHub's required human approvals or your existing CI checks.

A "safe to merge" verdict in our dashboard is guidance, not an approving GitHub review, and it only becomes a merge gate if your team has separately configured repository rules to read our check status.

The table below covers the signals you'll actually see and what each one calls for:

|  | Ito review signal | What it means | What to verify before merge |
| --- | --- | --- | --- |
|  | Failed run with a PR-attributable finding | At least one tested scenario related to the PR's changes failed. This is not proof of a configured GitHub merge block. | Reproduce or investigate the finding, check the latest commit's rerun, and obtain the required human approval. |
|  | Passed run | The scenarios generated from this PR's context completed without failure at the reviewed commit. | Confirm you're looking at the latest commit's results. Check independent GitHub approval or CI requirements. |
|  | Additional Findings | Pre-existing issues that do not set the run's pass/fail outcome. | Assess severity, assign ownership, and decide whether a Critical or High pre-existing issue warrants action before merge. |
|  | Review without usable execution evidence | Pending, Running, and Cancelled states, plus build failures and inaccessible services, are all distinct from a passed run. None is a clean behavioral pass. | Obtain current, usable execution evidence rather than inferring a clean result from an incomplete run. |

The one habit worth keeping regardless of the signal: always check the latest commit's completed run.

An old passed run from before your fix tells you nothing about whether the fix worked. Record ownership and rationale in your team's PR discussion or issue tracker, and if you want Ito to stop re-reporting a finding you've consciously accepted, turn on reviewer dismissal.

## Keep the tests that protect known behavior

None of this replaces a test suite. It changes what your suite is for.

Keep deterministic tests for the checks a runtime review shouldn't be responsible for, like security boundaries and data integrity rules. They also catch regressions on every push in milliseconds, while a runtime review takes 30 minutes to 2 hours. With automated review, you can get feedback on a PR before you've written any of those tests.

The DoltgreSQL example shows why both still matter. The t_inherit bug, where the named-window derivation failed to preserve its ordering, was a genuine coverage gap.

DoltHub's team wrote regression tests for it once the fix was confirmed, which is exactly how a suite is supposed to grow. The t_named bug is the harder case: DoltHub already had a test for that syntax, it just asserted the wrong result.

Neither bug is an argument against the suite. The suite tells you what the team intended; the runtime finding told them where that intention and the actual behavior had quietly diverged.

That's the honest limitation on both sides. A generated test plan only covers what its context pointed it toward. A source reviewer flags false positives and sometimes proposes the wrong fix.

A test suite can assert an incorrect result for years and pass every single CI run, the way t_named's did. None of that is a reason to drop any of the three signals, source review, runtime execution, deterministic tests. It's a reason not to trust any one of them past what it actually checked.

For a fuller view of how these fit together across your development process, see our post on [how to build quality into your software factory](https://www.ito.ai/blog/how-to-build-quality-into-your-software-factory).

## Where Ito fits in your GitHub review workflow

Ito's job is the middle step: turning intent into a plan, then a plan into evidence.

Your team still writes the PR description and any Test Instructions that state what matters. We translate that into scenarios, run them against the actual application, and hand back findings with reproduction evidence. Your engineers and your CI pipeline still own everything upstream and downstream of that.

That means the planning work doesn't disappear; it moves.

Writing a PR description that says what changed and how it should behave was already part of a good PR. Now, that description also drives what gets tested.

And because runtime evidence takes real setup and a real run to produce, it arrives after source review is already underway, not before. Read the diff first, let the run complete, then merge with both in hand.

The way to find out if that trade-off works for your team is a single-repository pilot: run it for a few weeks, and track which findings were reproducible, how much time and money you spent on environment setup, and whether the review time fits how your team actually merges. Judge it on findings that turned out to be real and verifiable, not on how many comments it left.

[Follow our quickstart guide to connect one GitHub repository and run a review on a representative pull request.](https://ito.ai/docs)