← All posts

7 min read

A Flutter Testing Strategy That Actually Pays Off

A practical Flutter testing strategy — unit, widget, golden and integration tests, mocking with mocktail, what not to test, and running it all in CI.

  • Flutter
  • Testing
  • Quality
Cover illustration for A Flutter Testing Strategy That Actually Pays Off

Most Flutter projects I’ve joined fall into one of two camps. Either there are no tests at all, or there’s a big folder of tests that nobody trusts, half of them skipped, and the rest breaking every time someone renames a widget.

Neither helps you ship. What you want is a small set of tests that catch real bugs, run fast, and only fail when something is actually wrong.

Testing is a big part of my day job right now — I own Playwright end-to-end testing at Cerebrum City — and the lessons carry straight over to Flutter. Here’s the strategy I use.

Think in layers, weighted by value

Flutter has three main kinds of tests, plus goldens as a special case of widget tests:

Type What it checks Speed How many
Unit Pure Dart logic: models, validators, repositories, state Milliseconds Lots
Widget One screen or component rendered in a test environment Fast A good number
Golden Pixel output of a widget against a reference image Fast-ish A few, for stable UI
Integration The real app on a device or emulator Slow A handful of key flows

The classic pyramid still holds: lots of cheap tests at the bottom, few expensive ones at the top. But I care less about the ratio and more about one question per test: would this catch a bug a user would notice?

Unit tests: where most of the value is

If your business logic lives in plain Dart classes — not inside widgets — it’s easy to test. This is one of the biggest practical payoffs of a sensible structure, which I cover in pragmatic clean architecture in Flutter.

Good unit-test targets:

  • Price, tax, and fee calculations
  • Validators and parsers
  • Date and time-zone logic (always more broken than you think)
  • State transitions in your Bloc, Cubit, or Notifier
  • Repository behavior when the network fails
void main() {
  group('BookingPrice', () {
    test('applies the platform fee on top of the base price', () {
      final price = BookingPrice(base: 5000, feePercent: 10);
      expect(price.total, 5500);
    });

    test('rejects negative base prices', () {
      expect(() => BookingPrice(base: -1, feePercent: 10), throwsArgumentError);
    });
  });
}

Fast, boring, and they catch exactly the kind of bug that ends up in a support ticket.

Mocks and fakes with mocktail

Unit tests need to swap out real dependencies like APIs and databases. I use mocktail because it needs no code generation:

class MockBookingApi extends Mock implements BookingApi {}

void main() {
  late MockBookingApi api;
  late BookingRepository repo;

  setUp(() {
    api = MockBookingApi();
    repo = BookingRepository(api);
  });

  test('returns cached bookings when the API is offline', () async {
    when(() => api.fetchBookings())
        .thenThrow(const SocketException('offline'));

    final result = await repo.loadBookings();

    expect(result.isFromCache, isTrue);
    verify(() => api.fetchBookings()).called(1);
  });
}

If you use any() as an argument matcher for a custom type, register a fallback first with registerFallbackValue(...) in setUpAll, or mocktail will complain.

Fakes vs mocks

Mocks are great for “throw this error” or “verify this was called”. But when a test sets up five when(...) calls just to get a dependency to behave sensibly, it’s usually a sign you want a fake instead: a small, real, in-memory implementation.

class FakeBookingStore implements BookingStore {
  final _items = <String, Booking>{};

  @override
  Future<void> save(Booking b) async => _items[b.id] = b;

  @override
  Future<Booking?> find(String id) async => _items[id];
}

Fakes behave like the real thing, are reusable across tests, and don’t break when you refactor which method calls which. My rule: mock the edges you need to control precisely, fake the stateful things.

Widget tests: screens without a device

Widget tests render a widget tree in a headless environment, let you tap and type, and check what’s on screen. They’re the best value for UI logic: does the error message appear, does the button disable while loading, does the empty state show.

testWidgets('shows an error when login fails', (tester) async {
  final auth = MockAuthService();
  when(() => auth.signIn(any(), any()))
      .thenThrow(InvalidCredentialsException());

  await tester.pumpWidget(MaterialApp(home: LoginScreen(auth: auth)));

  await tester.enterText(find.byKey(const Key('email')), 'a@b.com');
  await tester.enterText(find.byKey(const Key('password')), 'wrong');
  await tester.tap(find.text('Sign in'));
  await tester.pump(); // let the async call complete and rebuild

  expect(find.text('Email or password is incorrect'), findsOneWidget);
});

A few things that trip people up:

  • pumpWidget needs context. Most screens need a MaterialApp (for theme, directionality, navigation) and whatever providers your state management uses. Write a small pumpApp helper once and reuse it.
  • pump vs pumpAndSettle. pump() advances one frame. pumpAndSettle() keeps pumping until no frames are scheduled. That sounds convenient until a CircularProgressIndicator or any repeating animation is on screen — then pumpAndSettle never settles and the test times out. In those cases, use pump() with an explicit duration.
  • Prefer finders users would recognize. find.text('Sign in') reads like a spec. Use find.byKey where text is dynamic or localized.

Golden tests: use sparingly

A golden test renders a widget and compares it pixel by pixel against a stored PNG:

testWidgets('price card matches golden', (tester) async {
  await tester.pumpWidget(const MaterialApp(home: PriceCard(amount: 4999)));
  await expectLater(
    find.byType(PriceCard),
    matchesGoldenFile('goldens/price_card.png'),
  );
});

Generate or refresh the reference images with flutter test --update-goldens.

Goldens are excellent for design-system components that should never change by accident. They’re painful for screens that change every sprint. The other catch is platform differences: font rendering differs slightly between macOS, Linux and Windows, so a golden generated on your Mac may fail on a Linux CI runner. Pick one platform as the source of truth — usually CI — and generate goldens there, or load a fixed test font so text renders consistently.

Integration tests: only the flows that matter

The integration_test package runs your real app on a device or emulator, with real plugins:

void main() {
  IntegrationTestWidgetsFlutterBinding.ensureInitialized();

  testWidgets('user can book a session', (tester) async {
    app.main();
    await tester.pumpAndSettle();

    await tester.tap(find.text('Book'));
    await tester.pumpAndSettle();
    expect(find.text('Booking confirmed'), findsOneWidget);
  });
}

Run them with flutter test integration_test.

These are slow and need careful setup (test accounts, a stable backend or a mock server), so I keep them to the handful of journeys that would make the app worthless if broken: sign-up, login, the core action, and checkout. That’s exactly the mindset I describe in my Playwright end-to-end testing lessons — stable selectors, isolated test data, and ruthless focus on critical paths apply just as much on mobile.

What not to test

Some tests cost more than they’re worth:

  • Framework behavior. You don’t need a test that Text('Hello') shows “Hello”.
  • Trivial getters and data classes without logic.
  • Implementation details, like verifying a private method was called or asserting the exact widget tree structure. These tests break on every refactor and never on real bugs.
  • Third-party packages. Test your code’s reaction to them, via a mock or fake.

If a test regularly fails when nothing user-visible changed, fix it or delete it. A flaky test that everybody ignores is worse than no test.

Run it all in CI

Tests that only run on someone’s laptop eventually stop running. Put flutter analyze and flutter test on every pull request, and block merges on failures. Integration tests can run nightly or before releases if they’re too slow for every PR. I’ve written up a pipeline for this in Flutter CI/CD with Azure DevOps.

Coverage is a signal, not a goal

flutter test --coverage writes an lcov.info file you can visualize or upload to a coverage service. It’s useful for spotting important code that has no tests at all. It’s useless as a target: chasing a percentage produces tests that execute code without checking anything meaningful. I look at coverage per folder — is the pricing logic covered? the auth flow? — rather than one number for the whole app.

Takeaways

  • Put most effort into unit tests for business logic; keep that logic out of widgets so it’s testable.
  • Use mocktail for edges you need to control, and simple fakes for stateful dependencies.
  • Widget tests are the best value for UI behavior — and beware pumpAndSettle with endless animations.
  • Keep goldens for stable components and generate them on one consistent platform.
  • Limit integration tests to the few flows that would sink the app if broken.
  • Run tests on every PR, and treat coverage as a map of gaps, not a score.

The best test suite isn’t the biggest one — it’s the one your team trusts enough to ship on a Friday.

Comments

Questions, corrections or your own experience — leave a comment below (GitHub sign-in).