The SDET Playbook

← All questions

How should I manage test data: seeding, factories, synthetic data or production copies?

Asked Sep 28, 2026Viewed 0 times

1 Answer

Sign in to answer and to vote.

  • 0
    The SDET PlaybookSep 28, 2026

    Default to factories and small seeds, and avoid raw production data. A factory creates a valid object with sensible defaults so each test overrides only what it cares about, and updating the model means changing one place. Use scenario seeds for realistic states, per-test creation when isolation matters, and generated synthetic data (Faker-style libraries) when you need volume. If you must use production-like data, mask it in the copy pipeline before it reaches any non-production environment, because GDPR and similar rules make raw copies a liability. Start each CI run from a known clean state and clean up afterward. Vendor sources claim poor data causes a large share of failures; treat that figure as unverified.

    A factory can be as small as this:

    import { faker } from '@faker-js/faker';
    
    type User = { email: string; password: string; role: 'MEMBER' | 'ADMIN' };
    
    export const buildUser = (overrides: Partial<User> = {}): User => ({
      email: faker.internet.email(),
      password: faker.internet.password({ length: 12 }),
      role: 'MEMBER',
      ...overrides,
    });
    

    Create records through the API or a backend helper in a beforeEach hook, and remove them in afterEach or by rolling back a transaction, so each test leaves the environment as it found it.

    Sources: Faker docs, TestSprite, Keploy, QASkills