



Patrol and Maestro solve the same problem with opposite architectures: one runs inside your app's process, the other drives it from the outside. These are founding decisions, not feature gaps - no release will make them converge. Understand the trade-offs, and you can tell which tool fits your Flutter app before you look at any comparison table, including ours.
Patrol is an open-source agentic E2E testing framework for Flutter, built and maintained by us at LeanCode. Tests are written in Dart and run inside the app process, with direct access to the widget tree and to native automation: permission dialogs, notifications, connectivity toggles, GPS mocking, or WebViews. Tests compile to native Android and iOS test artifacts, so they run on local devices and simulators as well as any device farm that supports native tests.
Patrol 4.0 (December 2025) added support for Flutter Web and a VS Code extension. Patrol MCP (March 2026) connects AI agents to Patrol, letting them write, run, and debug E2E tests in a loop.
Maestro is an open-source mobile UI testing framework from mobile.dev. Tests are YAML flows executed by an external binary that drives the app through the accessibility tree - the same layer used by TalkBack and VoiceOver. That makes it framework-agnostic by design: native Android and iOS, React Native, Flutter, and Web (in beta) are all driven the same way.
Maestro 2.2 (February 2026) added visual regression testing with the assertScreenshot command. Maestro 2.6 (May 2026) introduced Maestro Viewer, which embeds a live device inside an AI coding agent, shipping alongside Maestro MCP. Flows can also be recorded visually in Maestro Studio and run on hosted infrastructure in Maestro Cloud.
| Patrol | Maestro |
|---|---|
Test language | |
| Dart | YAML (+ JavaScript for logic) |
Architecture | |
| Inside the app process (gray box) | External binary via accessibility tree (black box) |
Supported frameworks | |
| Flutter: Android, iOS, Web, macOS (alpha) | Flutter, native Android/iOS, React Native, Web (beta) |
Selectors | |
| Keys, widget types, any widget property | Visible text, accessibility labels |
Physical iOS devices locally | |
| Yes | No - device farms or unofficial patches |
Where tests run | |
| Local devices and simulators; any farm that runs native tests, e.g. BrowserStack, Firebase Test Lab, Marathon Cloud | Android devices and simulators, iOS simulators; Maestro Cloud; farms with a Maestro integration, e.g. BrowserStack, TestingBot, LambdaTest |
App restart testing | |
| No | Yes |
Network and location control | |
| Airplane mode, Wi-Fi, cellular, Bluetooth, GPS (iOS network toggles: physical devices) | Airplane mode (Android only), GPS |
AI tooling | |
| Patrol MCP + agent skills | Maestro MCP, Maestro Viewer |
State access and mocking from the test | |
| Yes: app state, cubits, injected mocks | No |
Code reuse | |
| Functions, types, IDE refactoring | YAML subflows + JavaScript |
Setup | |
| Project integration required | CLI install, no project changes |
Open source | |
| Yes | Yes (Cloud is paid) |
Patrol tests are Dart code running inside your app's process. They see what the app sees: the widget tree, the state, every mock you put in the code. Maestro is a binary that drives the app from the outside and reads the accessibility tree.
Neither approach is "more correct". It's a trade. The black box buys independence from the application - it tests whatever binary you hand it, survives an app restart, and works with any framework - and pays with no access to the app's internals. The gray box buys the inside of the app - the widget tree, the state, the mocks you build in - and pays by living inside its process.
Everything below is a consequence of this trade.
Here is the same login in both tools:
// Patrol
await $.pumpWidgetAndSettle(const MyApp());
await $(#email).enterText('user@leancode.co');
await $(#password).enterText('secret');
await $(#loginButton).tap();
await $(#homeScreen).waitUntilVisible();# Maestro
- launchApp
- tapOn: "Email"
- inputText: "user@leancode.co"
- tapOn: "Password"
- inputText: "secret"
- tapOn: "Log in"
- assertVisible: "Welcome"Both are readable, and at five tests either one will do. The difference shows up later: more people writing tests, the same steps repeating in every one of them, and someone has to keep it all consistent. Teams usually start feeling it around 15 flows. Past 30, keeping them clean becomes a project of its own.
The difference is reuse. In Patrol, a login is a function - with types, IDE navigation, and refactoring:
Future<void> logIn(
PatrolIntegrationTester $, {
required String email,
required String password,
}) async {
await $(#email).enterText(email);
await $(#password).enterText(password);
await $(#loginButton).tap();
await $(#homeScreen).waitUntilVisible();
}
// in the test
await logIn($, email: 'user@leancode.co', password: 'secret');In Maestro, reuse means a second file. The helper lives on its own and the test calls it through runFlow, passing parameters as environment variables:
# subflows/login.yaml - the helper itself
- tapOn: "Email"
- inputText: ${EMAIL}
- tapOn: "Password"
- inputText: ${PASSWORD}
- tapOn: "Log in"# login_test.yaml - the test that uses the helper
- runFlow:
file: subflows/login.yaml
env:
EMAIL: "user@leancode.co"
PASSWORD: "secret"A Maestro flow is a list of steps, and that is what makes the first one so easy to read. Anything the list can't express - a value computed at runtime, a call to your API - goes to JavaScript, which Maestro added as the way out. Here is the same login with a user created through your API:
// scripts/create_user.js
const res = http.post('https://api.example.com/test-users', {
body: JSON.stringify({ plan: 'premium' }),
});
output.user = json(res.body);# subflows/login.yaml
- runScript: ../scripts/create_user.js
- tapOn: "Email"
- inputText: ${output.user.email}In Patrol, creating that user is Dart calling Dart: one language, typed values, and as many login variants as you need sitting in one helper file - you change the call, not the helper. In Maestro, the same step spans two languages, with values crossing between them as untyped strings. The suite also fragments into a directory of micro-files, and Maestro's own best practices recommend keeping helpers in a common/ folder and excluding it in config.yaml so they don't run as standalone tests. Paths are resolved inconsistently, so reorganizing directories breaks them (#2792), and paths in runFlow can't be parameterized at all (#1314). None of it is unworkable. It is just work that Dart does for you.
Maestro's own comparison with Appium puts it plainly: "Use Appium for complex, custom workflows requiring advanced programming. Choose Maestro for fast, reliable, and accessible testing, especially for React Native projects." It is a fair split, and it maps onto time: accessible while the suite is small, programmable once it carries real business flows. For Flutter, the programmable one is Patrol.
When a test fails without a real bug behind it, the selector is usually the reason - backend and system instability aside, which no test framework controls. In Flutter apps, that is the main reason Patrol suites go red less often. Text-based selectors work, but every copy change and every new language turns into test maintenance, and that cost grows with the suite.
Patrol selects widgets by whatever your code already has: Key values, widget types, text, or any widget property. Keys are recommended, not required, and they survive copy changes and localization.
Maestro reads the semantics tree, and its documentation states that Flutter Keys can't be used there. In React Native, it would simply read a testID, a plain property on the component. Flutter has no equivalent: identity has to be added to the accessibility tree itself - wrap the widget in Semantics(identifier: 'signin_button') and target it with id. That works, but the accessibility tree isn't there just for your tests. Standard accessibility work like merging or excluding semantics can quietly move an identifier onto a different element, or remove it: we have seen a tap by id land on the wrong button while the step still reported success. Skip the annotations and you are back to matching on text that changes with every copy refresh. A Key never leaves the widget tree, so it survives all of it. Of all the frameworks Maestro can drive, Flutter is the one that asks the most in return.
One more note on stability comparisons between tools: the cheapest way to stabilize a suite is to test less, and a suite that tests less is not the goal. Permission dialogs are a good example. Maestro grants all permissions by default at launch, so its flows never meet one. Patrol runs the app the way your users get it, dialogs included - and once you have covered that flow, you can skip the prompts in the tests that follow. A suite that never opens that door will never fail at it. Compare stability at equal coverage, or the numbers don't mean much.
AI is where both tools are moving fastest. Maestro MCP ships with the CLI and lets coding agents - Claude Code, Cursor, Copilot and others - inspect the screen, run flows, and take screenshots, with Maestro Viewer embedding a live device directly inside the agent. On top of that, experimental AI commands like assertWithAI validate visual states from a natural-language description, for cases where element-based assertions fall short.
Patrol ships Patrol MCP - the same set of agents, Claude Code, Cursor, Copilot and others - together with agent rules: coding best practices and, optionally, a test architecture for the agent to follow. The result is an agent that writes solid, readable tests from day one - and debugs them as it goes, because the MCP runs them in a loop.
Both tools plug into an agent, so the real question is how much that agent has to work with. An agent driving Maestro works with what is currently on screen: it discovers a flow one screen at a time, reactively, unless you describe the flow to it upfront. An agent driving Patrol has the app's source code next to the test, so it reads whole flows, their states, and their edge cases before it runs anything. And because tests are code, it reuses the helpers your team already wrote - past the thirty flows where YAML starts to strain, a new test is mostly a composition of steps that already exist, so each one costs less than the last.
That context matters most when a test goes red. A black-box agent only sees that something expected isn't on screen and has to guess whether the app broke or the test is out of date; its options are a retry, a different text selector, or a longer wait. An agent working with Patrol compares the test against the code, the widget state, and what is on screen, so it can tell a real regression from a renamed key or a copy change: fix the test in one case, report the bug in the other. That is the difference between an agent that executes tests and an agent that explores your app, finds actual bugs, and heals tests when they break. On Flutter, that context exists only inside the app's process - which is where Patrol runs.
On Android, Patrol runs on emulators and physical devices, against debug or release builds, with the same native automation everywhere. Maestro covers Android well too, driving emulators and physical devices alike.
The differences start on iOS. Patrol runs on simulators and on physical iPhones, like any Flutter app. Maestro runs locally on iOS simulators only, and simulators don't expose OS-level features like connectivity toggles or camera-based QR scanning - those scenarios stay out of reach.
The bigger difference is controlling the device from a test. Patrol toggles Wi-Fi, cellular, and Bluetooth individually on both platforms (on iOS, on physical devices), while Maestro offers airplane mode on Android only. If your user journeys cross the boundary between Flutter and the operating system, and in production apps they do, Patrol is the tool that covers them end-to-end on both platforms.
On the web, both run tests in Chromium: Patrol against Flutter Web, and Maestro with web support that is still in beta.
Enterprise teams have their own CI and their own infrastructure requirements. The tool has to fit in, not the other way around.
Patrol tests compile to native test artifacts (Android instrumentation tests and XCUITest), so they run anywhere native tests run - Firebase Test Lab, AWS Device Farm, BrowserStack and Marathon Cloud, among them - or in your own device lab. On the web, the same tests run through Playwright, headless in CI. Sharding splits the suite across your devices, so a farm run scales with the devices you pay for. No vendor lock-in on infrastructure.
Maestro runs locally on simulators and emulators. Physical iOS devices have been the most requested feature since January 2023 (#686) - the Maestro team has signaled work toward official support, but as of publication it hasn't shipped. Maestro Cloud doesn't close that gap either: its docs run iOS tests on simulators and tell you not to upload device builds. Hosted runs start at $250 per device monthly. Real iPhones today mean a third-party farm that has built its own Maestro integration - BrowserStack, TestingBot, LambdaTest - or unofficial community tooling.
After every step, Maestro waits for the UI to settle before moving on. That auto-wait is a feature - it absorbs flakiness without manual waits - but it's paid on every step of every test. Patrol reacts to the widget state from inside the process, so a step ends the moment the app is ready, not when the screen looks quiet. The difference compounds with the size of the suite.
Scrolling makes this visible. Both tools scroll the same way in principle: move by a step, check for the target, repeat. The difference is what a "check" costs. Patrol checks the widget tree in-process, between frames, and drives the actual scrollable it found in the tree. Maestro swipes, then reads the accessibility hierarchy from the outside and waits for the UI to settle before the next step. Same loop, very different price per iteration.
// Patrol - re-checks in process and keeps moving
await $(keys.agenda.sessionTile(title)).scrollTo().tap();# Maestro - swipes, then dumps the accessibility hierarchy and waits
- scrollUntilVisible:
element:
text: ".*Testing Strategies for Flutter Applications.*"
- tapOn: ".*Testing Strategies for Flutter Applications.*"Maestro at stock defaults; Patrol's scroll step matched to a default Maestro swipe, so both cover the same ground per iteration.
And it's not just scrolling. Here is the same form - thirteen interactions, from a text field to a color picker - filled by both tools:
Every real project has blockers, and this is where the gray box adapts to your setup instead of the other way around.
Your app talks to hardware? Patrol drives a mock that lives in the app code and controls it at runtime from the test. A black box needs a dedicated build with the mock frozen in, a separate binary to maintain, and no control over it during the test.
Your flow needs a human in the loop - an account activation or a credit approval that a back-office admin has to confirm? A black-box tool can click and, at best, send a request, so the flow stalls. A Patrol test can do the same: call your backend to approve it, mock the dependency, or set the state directly. The test gets unblocked.
The same control removes overlap from the suite: instead of every test clicking through registration and onboarding, one test covers that path, and the rest start already signed in: the test puts the session for whatever account the case calls for straight into the app's state, from code. And when tapping and scrolling aren't enough, a Patrol test can read a cubit or any widget property, including its color. You can go that low - you just rarely have to. When you do, there's a way in.
Setup is Maestro's strongest card. The CLI installs with a one-line script, requires no changes in your project, and the first flow runs within an hour. For a non-technical QA team, that's a real difference.
Patrol asks for more upfront: integration with your project and its native dependencies, with the most friction on iOS. That's the price of the gray box, and it's a one-time cost. A minimal Android setup, though, is quick - there's even a ready agent skill that takes a coding agent from zero to a first green test. Enough to get a genuine feel for Patrol before committing to full integration.
For enterprise Flutter projects, scale is the headline: the suite will grow, and maintainability decides its cost more than any other factor. Beyond that, the checklist we see teams evaluate against:
There is no ranking to close with, because the architecture decides, not a feature list. Three signals map your project to the right column: the size and expected lifetime of the test suite, the diversity of your stack, and who will write the tests.
You know the trade-offs now - independence from the app versus access to its inside. That answer is specific to your project, and you already have everything needed to make it.
If you're going the Patrol route, we can help with Patrol Setup & Training and Automated UI Testing in Flutter. And if you want to see what AI agents do with it, read about Patrol MCP.

PATROL / TESTING
12 minJul 7, 2026Patrol has evolved far beyond being just a Flutter testing tool. From Android, iOS, and Web support to integrating with cloud device farms, native automation, and AI-powered testing, see how it’s redefining Flutter E2E testing, no matter if you’ve been following Patrol for a while or you’re new to it.

PATROL / TESTING
10 minDec 8, 2025Patrol 4.0 is here! By far the biggest update since improvements were made to the test building, it brings support for a web platform, a VS Code extension, a better debugging experience, and many smaller improvements. Let’s dive into the details and the brief backstory behind this release.

PATROL / TESTING
10 minDec 8, 2025Patrol has reached version 4.0, marking a major milestone. Among the many new features, one stands out in particular: Patrol now supports Web! In this article, you’ll find a rundown of what Patrol Web can do, but also a look behind the scenes: how we designed it, why certain decisions were made, and what it took to bring Patrol’s architecture from mobile into the browser.