•7 min read

Database Seeding Strategies with Prisma and TypeScript

Database Seeding Strategies with Prisma and TypeScript

Database seeding is often treated as an afterthought—a quick script thrown together to dump a few rows of mock data into a development environment. However, as applications scale and domain models grow increasingly complex, a brittle seed script quickly becomes a significant bottleneck for developer productivity and automated testing.

In the modern Node.js ecosystem, combining Prisma and TypeScript provides an incredibly powerful foundation for building robust, type-safe database seeding tools. In this guide, we will explore advanced strategies for database seeding that go far beyond the typical tutorial. We will focus on programmatic seeding, integrating Faker.js for realistic mock data, and, crucially, how to elegantly handle deeply nested relational data.

Audio Briefing
0:00 / 0:00

The Foundation: Why Basic Seeding Fails

The typical introductory Prisma tutorial suggests writing a simple script that executes a series of prisma.user.create() calls. While this works for day one, it falls apart by day ten. Basic scripts suffer from several critical flaws:

  1. Lack of Idempotency: Running the script twice results in unique constraint violations.
  2. Hardcoded Data: Statically defined data limits the ability to test pagination, search features, or complex UI states.
  3. Relational Spaghetti: Managing foreign keys manually across dozens of tables is error-prone.
  4. Performance Issues: Awaiting thousands of individual insert operations sequentially will drastically slow down CI pipelines.

To build a professional-grade seeding strategy, we need to address each of these challenges systematically.

Advertisement

Idempotency and Database Cleansing

A seed script must be idempotent. You should be able to run it repeatedly without breaking the database state or encountering errors. There are generally two approaches to achieve this: aggressive cleansing or careful upserting.

Strategy 1: Aggressive Cleansing

If you want a completely fresh state every time you seed, you need to wipe the database. With Prisma, deleting records in the correct order to respect foreign key constraints can be tedious. Instead of manually deleting from each table, you can dynamically truncate tables.

import { PrismaClient } from '@prisma/client';

const prisma = new PrismaClient();

async function cleanDatabase() {
  const tableNames = await prisma.$queryRaw<
    Array<{ tablename: string }>
  >`SELECT tablename FROM pg_tables WHERE schemaname='public'`;

  const tables = tableNames
    .map(({ tablename }) => tablename)
    .filter((name) => name !== '_prisma_migrations')
    .map((name) => `"public"."${name}"`)
    .join(', ');

  try {
    await prisma.$executeRawUnsafe(`TRUNCATE TABLE ${tables} CASCADE;`);
    console.log('Database cleaned successfully.');
  } catch (error) {
    console.error('Error cleaning database', error);
  }
}

Note: The snippet above uses PostgreSQL syntax. The approach will vary depending on your underlying database engine.

Strategy 2: Upserting for Core Data

For core configurations, roles, or taxonomies that must exist, upsert is your best friend. upsert guarantees that a record exists without throwing an error if it is already present.

async function seedRoles() {
  const roles = ['ADMIN', 'USER', 'EDITOR'];

  for (const roleName of roles) {
    await prisma.role.upsert({
      where: { name: roleName },
      update: {},
      create: { name: roleName },
    });
  }
}

By mixing truncation for mock data and upserts for foundational data, you create a robust, repeatable seeding baseline.

Integrating Faker.js for Realistic Data

To simulate real-world usage, we need a substantial volume of realistic data. @faker-js/faker is the industry standard for this.

When integrating Faker into your seed scripts, the most important (and frequently overlooked) step is setting a deterministic seed. This ensures that every developer on your team, and your CI environment, generates the exact same "random" data.

import { faker } from '@faker-js/faker';

// Set a deterministic seed for consistent generation
faker.seed(12345);

function createUserFactory() {
  const firstName = faker.person.firstName();
  const lastName = faker.person.lastName();
  
  return {
    email: faker.internet.email({ firstName, lastName }).toLowerCase(),
    name: `${firstName} ${lastName}`,
    bio: faker.lorem.paragraph(),
    avatarUrl: faker.image.avatar(),
  };
}

Using factory functions keeps your main seeding logic clean and makes it trivial to generate arrays of data.

Mastering Relational Data

The true complexity of database seeding lies in managing relationships. Prisma's nested writes provide an elegant solution for creating graphs of related data in a single transaction.

Deep Writes

Instead of creating a User, fetching their ID, and then creating Posts mapped to that ID, you can do it all at once:

async function seedUserWithPosts() {
  await prisma.user.create({
    data: {
      ...createUserFactory(),
      posts: {
        create: Array.from({ length: 5 }).map(() => ({
          title: faker.lorem.sentence(),
          content: faker.lorem.paragraphs(3),
          published: faker.datatype.boolean(),
        })),
      },
    },
  });
}

This approach is highly readable and ensures referential integrity without manual ID tracking.

Resolving Complex Relationships

What if you have many-to-many relationships, or relationships to records that were generated dynamically? For example, assigning random Tags to Posts.

To handle this, you need to first generate the pool of related entities, and then randomly select from them during the creation phase.

async function seedComplexGraph() {
  // 1. Create a pool of tags
  const tagsData = Array.from({ length: 10 }).map(() => ({
    name: faker.word.noun(),
  }));
  
  await prisma.tag.createMany({ data: tagsData });
  const allTags = await prisma.tag.findMany();

  // 2. Helper to get random tags
  const getRandomTags = (count: number) => {
    return faker.helpers.arrayElements(allTags, count).map(tag => ({
      id: tag.id
    }));
  };

  // 3. Create posts and connect random tags
  await prisma.post.create({
    data: {
      title: faker.lorem.sentence(),
      content: faker.lorem.paragraphs(),
      tags: {
        connect: getRandomTags(3),
      },
      author: {
        create: createUserFactory(),
      }
    }
  });
}

Using the connect syntax allows you to link newly created records to existing records effortlessly.

Advertisement

Performance: Batching and Transactions

When scaling your seed script to generate tens of thousands of records, awaiting individual prisma.model.create() calls will result in severe performance degradation due to network overhead and database roundtrips.

To optimize, utilize createMany combined with chunking.

async function seedLargeVolumeOfUsers() {
  const totalUsers = 10000;
  const batchSize = 1000;
  
  for (let i = 0; i < totalUsers; i += batchSize) {
    const usersBatch = Array.from({ length: batchSize }).map(createUserFactory);
    
    await prisma.user.createMany({
      data: usersBatch,
      skipDuplicates: true,
    });
    
    console.log(`Seeded batch ${i / batchSize + 1}`);
  }
}

createMany executes a single INSERT statement for the entire array, speeding up the process by orders of magnitude.

If you have complex relational data that cannot use createMany (as createMany does not support nested relations), you can fall back to Prisma transactions (prisma.$transaction) to group multiple create calls into a single database transaction, significantly reducing commit overhead.

Structuring the Seed Directory

As your seed logic grows, a single seed.ts file becomes unmanageable. Adopt a factory and runner pattern:

prisma/
  seed/
    index.ts        # The main entry point orchestrator
    factories/
      user.ts       # User factory functions
      post.ts       # Post factory functions
    runners/
      roles.ts      # Logic to seed foundational roles
      mockData.ts   # Logic to orchestrate Faker data

Your index.ts simply coordinates the runners:

import { PrismaClient } from '@prisma/client';
import { seedRoles } from './runners/roles';
import { seedMockData } from './runners/mockData';
import { cleanDatabase } from './utils/clean';

const prisma = new PrismaClient();

async function main() {
  console.log('Starting seed process...');
  
  if (process.env.NODE_ENV !== 'production') {
    await cleanDatabase();
  }
  
  await seedRoles(prisma);
  
  if (process.env.NODE_ENV === 'development') {
    await seedMockData(prisma);
  }
  
  console.log('Seed completed.');
}

main()
  .catch((e) => {
    console.error(e);
    process.exit(1);
  })
  .finally(async () => {
    await prisma.$disconnect();
  });

Conclusion

Database seeding is not just a chore; it is a critical component of a robust engineering culture. By leveraging TypeScript's type safety, Prisma's expressive ORM syntax, and Faker.js, you can build seeding infrastructure that empowers your team rather than holding them back.

Embrace idempotency, utilize nested writes for relationships, and optimize with batching. Your future self—and your QA team—will thank you.

You Might Also Like

Share this article:

Stay Updated

Get the latest posts delivered straight to your inbox.

Free Developer Utilities

Free In-Browser Developer Tools

Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.

Explore Tools
Advertisement