How should I manage test data: seeding, factories, synthetic data or production copies?
Asked by The SDET Playbook
Asked Sep 28, 2026Viewed 0 times
How should I manage test data: seeding, factories, synthetic data or production copies?
Asked by The SDET Playbook
Sign in to answer and to vote.
Default to factories and small seeds, and avoid raw production data. A factory creates a valid object with sensible defaults so each test overrides only what it cares about, and updating the model means changing one place. Use scenario seeds for realistic states, per-test creation when isolation matters, and generated synthetic data (Faker-style libraries) when you need volume. If you must use production-like data, mask it in the copy pipeline before it reaches any non-production environment, because GDPR and similar rules make raw copies a liability. Start each CI run from a known clean state and clean up afterward. Vendor sources claim poor data causes a large share of failures; treat that figure as unverified.
A factory can be as small as this:
import { faker } from '@faker-js/faker';
type User = { email: string; password: string; role: 'MEMBER' | 'ADMIN' };
export const buildUser = (overrides: Partial<User> = {}): User => ({
email: faker.internet.email(),
password: faker.internet.password({ length: 12 }),
role: 'MEMBER',
...overrides,
});
Create records through the API or a backend helper in a beforeEach hook, and remove them in afterEach or by rolling back a transaction, so each test leaves the environment as it found it.
Sources: Faker docs, TestSprite, Keploy, QASkills