•13 min read

Turborepo Remote Caching at Scale: Pruned Docker Builds, GCS Cache & CI Pipeline Speedup

Turborepo Remote Caching at Scale: Pruned Docker Builds, GCS Cache & CI Pipeline Speedup

Turborepo is a high-performance build system for JavaScript and TypeScript monorepos. Its core value proposition lies in intelligent task scheduling, content-addressable caching, and remote caching. This guide details the implementation of a robust Turborepo setup, focusing on optimized Docker builds, GCS-backed remote caching, and significant CI pipeline acceleration.

Audio Briefing
0:00 / 0:00

Understanding Turborepo's Caching Mechanism

Turborepo's caching operates on a content-addressable principle. For any given task (e.g., build, test, lint), Turborepo computes a hash based on:

  1. Input files: Source code, configuration files, and assets configured within turbo.json under the inputs key.
  2. Dependencies: package.json files of the current package and its transitive dependencies.
  3. Environment variables: Runtime variables explicitly declared within turbo.json under the env key.
  4. Turborepo configuration: Global settings inside the root turbo.json file.
  5. Task command: The exact command being executed.

If a task with an identical hash has been executed previously, Turborepo retrieves the cached output instead of re-running the task. This output can be stored locally or in a remote cache.

Advertisement

Architecture Overview

Our optimized architecture integrates Turborepo with Docker and Google Cloud Storage (GCS) for a scalable, efficient build process:

  1. Turborepo Monorepo: Centralized codebase for multiple applications and packages.
  2. turbo prune: Generates a minimal subset of the monorepo required to build a specific application, optimizing Docker build contexts.
  3. Multi-stage Docker Builds: Leverages turbo prune output and Docker's layer caching for lean, reproducible images.
  4. GCS Remote Cache: A shared, persistent cache backend for Turborepo, accessible by all CI agents and developers.
  5. GitHub Actions: Orchestrates the build, test, and deployment workflows, utilizing the remote cache.

Architectural Tradeoffs

FeatureProsCons
turbo pruneMinimal Docker context, faster builds, smaller imagesAdds complexity to Dockerfile, requires careful turbo.json inputs definition
Multi-stage DockerSmaller final images, better layer cachingMore complex Dockerfiles, potential for build-time dependency bloat if not careful
GCS Remote CacheShared cache, high availability, scalability, reduced CI timeCost of GCS, initial setup complexity, network latency for cache misses
GitHub ActionsIntegrated CI/CD, good ecosystemVendor lock-in, potential for rate limits on large repos

Implementing turbo prune for Optimized Docker Builds

turbo prune is critical for Docker builds. It creates a new directory containing only the packages and their dependencies required to build a specific target application within the monorepo. This drastically reduces the Docker build context size, leading to faster image builds and smaller images.

Consider a monorepo structure:

/
├── apps/
│   ├── web/
│   └── api/
├── packages/
│   ├── ui/
│   ├── utils/
│   └── db/
├── turbo.json
├── package.json
└── pnpm-lock.yaml

To build apps/api, we only need apps/api, packages/utils, and packages/db. turbo prune isolates this.

turbo.json Configuration

Ensure your turbo.json accurately defines inputs and outputs for tasks. This is crucial for correct hashing and pruning.

{
  "$schema": "https://turbo.build/schema.json",
  "pipeline": {
    "build": {
      "dependsOn": ["^build"],
      "outputs": ["dist/**", ".next/**", "build/**"],
      "inputs": ["src/**/*.ts", "src/**/*.tsx", "src/**/*.js", "src/**/*.jsx", "tsconfig.json", "package.json"]
    },
    "test": {
      "dependsOn": ["^build"],
      "outputs": [],
      "inputs": ["src/**/*.test.ts", "src/**/*.test.tsx", "src/**/*.spec.ts", "src/**/*.spec.tsx"]
    },
    "lint": {
      "outputs": []
    },
    "dev": {
      "cache": false,
      "persistent": true
    }
  }
}

Dockerfile with turbo prune

Here's a multi-stage Dockerfile for an api application, leveraging turbo prune.

# Stage 1: Build dependencies and application
FROM node:20-alpine AS builder

# Install pnpm globally
RUN corepack enable && corepack prepare pnpm@latest --activate

# Set working directory
WORKDIR /app

# Copy lockfile and package.json files first to leverage Docker cache
# This layer changes infrequently, maximizing cache hits
COPY pnpm-lock.yaml ./
COPY package.json ./
COPY turbo.json ./

# Copy only the pruned monorepo subset for the 'api' application
# This command is executed by the CI/CD pipeline or locally before `docker build`
# Example: `turbo prune --scope=api --docker`
# The output of `turbo prune` is expected in the './.pruned-build' directory
COPY .pruned-build/pnpm-workspace.yaml ./.pruned-build/
COPY .pruned-build/full/ ./.pruned-build/full/
COPY .pruned-build/json/ ./.pruned-build/json/

# Install dependencies using pnpm
# Use --frozen-lockfile to ensure reproducible builds
RUN pnpm install --frozen-lockfile --prod=false

# Copy the actual source code for the pruned packages
# This is crucial: `turbo prune` only copies package.json and lockfiles,
# not the source code itself. We copy it from the original context.
# The `full` directory contains symlinks to the original source.
# We need to copy the actual files.
# This assumes the Docker build context is the root of the monorepo.
COPY apps/api ./apps/api
COPY packages/db ./packages/db
COPY packages/utils ./packages/utils

# Build the 'api' application
# Turborepo will use its cache or build if necessary
RUN pnpm turbo run build --filter=api...

# Stage 2: Production image
FROM node:20-alpine AS runner

# Install pnpm globally (for production dependencies if needed, though often not)
RUN corepack enable && corepack prepare pnpm@latest --activate

WORKDIR /app

# Copy only production dependencies from the builder stage
# This ensures a minimal production image
COPY --from=builder /app/pnpm-lock.yaml ./
COPY --from=builder /app/package.json ./
COPY --from=builder /app/turbo.json ./
COPY --from=builder /app/.pruned-build/pnpm-workspace.yaml ./.pruned-build/
COPY --from=builder /app/.pruned-build/json/ ./.pruned-build/json/

# Install production dependencies
RUN pnpm install --prod --frozen-lockfile

# Copy the built application artifacts from the builder stage
# This includes `dist` directories for the API and its dependencies
COPY --from=builder /app/apps/api/dist ./apps/api/dist
COPY --from=builder /app/packages/db/dist ./packages/db/dist
COPY --from=builder /app/packages/utils/dist ./packages/utils/dist

# Expose port (if applicable)
EXPOSE 3000

# Define the command to run the application
CMD ["node", "apps/api/dist/index.js"]

Explanation of turbo prune in Docker context:

  1. turbo prune --scope=api --docker: This command, executed before docker build, creates a .pruned-build directory at the monorepo root.
    • pnpm-workspace.yaml: Copied to .pruned-build/pnpm-workspace.yaml.
    • full/: Contains symlinks to package.json files of all relevant packages.
    • json/: Contains package.json files of all relevant packages, but flattened.
  2. COPY .pruned-build/...: We copy these generated files into the Docker image.
  3. pnpm install: With the pruned package.json files and pnpm-lock.yaml, pnpm installs only the necessary dependencies.
  4. COPY apps/api ./apps/api: Crucially, turbo prune does not copy source files. You must explicitly copy the source code for the target application and its direct dependencies from the original monorepo context. This is why the docker build command should be run from the monorepo root.
  5. pnpm turbo run build --filter=api...: Turborepo builds the api application and its upstream dependencies.

Self-Hosted Remote Cache with GCS

Turborepo supports various remote cache backends. For enterprise environments, a self-hosted solution using cloud storage (GCS, S3) offers control, scalability, and cost-effectiveness.

GCS Bucket Setup

  1. Create a GCS bucket: gs://your-turborepo-cache-bucket
  2. Ensure appropriate IAM permissions: The service account or user accessing the cache needs Storage Object Admin or Storage Object Creator and Storage Object Viewer roles.

Turborepo Configuration for GCS

Turborepo uses environment variables to configure remote caching.

# .env.local or CI environment variables
TURBO_REMOTE_CACHE_SIGNATURE_KEY="your_strong_secret_key_for_signing_cache_requests"
TURBO_REMOTE_CACHE_READ_ONLY="false" # Set to true for read-only environments
TURBO_REMOTE_CACHE_GCS_BUCKET="your-turborepo-cache-bucket"
TURBO_REMOTE_CACHE_GCS_SERVICE_ACCOUNT_KEY="path/to/your/gcs-service-account-key.json" # Or base64 encoded JSON

Important Security Note: Never commit TURBO_REMOTE_CACHE_SIGNATURE_KEY or TURBO_REMOTE_CACHE_GCS_SERVICE_ACCOUNT_KEY directly into your repository. Use environment variables, secret management systems (e.g., GitHub Secrets, GCP Secret Manager).

For CI/CD, you'll typically base64 encode the service account key JSON and store it as a secret.

# Example for GitHub Actions
echo "${{ secrets.GCP_SERVICE_ACCOUNT_KEY_BASE64 }}" | base64 --decode > gcs-key.json
export TURBO_REMOTE_CACHE_GCS_SERVICE_ACCOUNT_KEY=$(pwd)/gcs-key.json
Advertisement

CI Pipeline Speedup with GitHub Actions

Integrating the above components into a GitHub Actions workflow dramatically reduces CI build times.

name: CI Build & Test

on:
  push:
    branches:
      - main
  pull_request:
    branches:
      - main

jobs:
  build-api:
    runs-on: ubuntu-latest
    steps:
      - name: Checkout repository
        uses: actions/checkout@v4
        with:
          fetch-depth: 0 # Required for Turborepo to correctly detect changed files

      - name: Setup Node.js
        uses: actions/setup-node@v4
        with:
          node-version: 20
          cache: 'pnpm' # Cache pnpm dependencies

      - name: Install pnpm
        run: corepack enable && corepack prepare pnpm@latest --activate

      - name: Install monorepo dependencies
        run: pnpm install --frozen-lockfile

      - name: Configure Turborepo Remote Cache (GCS)
        env:
          TURBO_REMOTE_CACHE_SIGNATURE_KEY: ${{ secrets.TURBO_REMOTE_CACHE_SIGNATURE_KEY }}
          TURBO_REMOTE_CACHE_GCS_BUCKET: your-turborepo-cache-bucket
          GCP_SERVICE_ACCOUNT_KEY_BASE64: ${{ secrets.GCP_SERVICE_ACCOUNT_KEY_BASE64 }}
        run: |
          echo "${{ env.GCP_SERVICE_ACCOUNT_KEY_BASE64 }}" | base64 --decode > gcs-key.json
          export TURBO_REMOTE_CACHE_GCS_SERVICE_ACCOUNT_KEY=$(pwd)/gcs-key.json
          echo "TURBO_REMOTE_CACHE_GCS_SERVICE_ACCOUNT_KEY=$(pwd)/gcs-key.json" >> $GITHUB_ENV
          echo "TURBO_REMOTE_CACHE_SIGNATURE_KEY=${{ secrets.TURBO_REMOTE_CACHE_SIGNATURE_KEY }}" >> $GITHUB_ENV
          echo "TURBO_REMOTE_CACHE_GCS_BUCKET=your-turborepo-cache-bucket" >> $GITHUB_ENV

      - name: Build API application with Turborepo
        run: pnpm turbo run build --filter=api...

      - name: Run API tests with Turborepo
        run: pnpm turbo run test --filter=api...

      - name: Lint API application with Turborepo
        run: pnpm turbo run lint --filter=api...

      - name: Prune API for Docker build
        run: pnpm turbo prune --scope=api --docker

      - name: Build Docker image for API
        run: docker build -t your-registry/api:${{ github.sha }} -f apps/api/Dockerfile .

      - name: Push Docker image (example)
        # Add Docker login steps here if pushing to a private registry
        # run: docker push your-registry/api:${{ github.sha }}
        run: echo "Docker image built: your-registry/api:${{ github.sha }}"

Key optimizations in the CI workflow:

  • actions/checkout@v4 with fetch-depth: 0: Ensures Turborepo has access to the full Git history for accurate change detection and caching.
  • actions/setup-node@v4 with cache: 'pnpm': Caches node_modules for pnpm, reducing pnpm install time on subsequent runs.
  • Turborepo Remote Cache Configuration: Environment variables are set up to point Turborepo to the GCS cache.
  • pnpm turbo run build --filter=api...: Turborepo intelligently builds only the api application and its dependencies. If a cache hit occurs, this step completes almost instantly.
  • pnpm turbo prune --scope=api --docker: Creates the optimized build context for Docker.
  • docker build ... -f apps/api/Dockerfile .: Builds the Docker image. Because turbo prune reduced the context and the Dockerfile is multi-stage, this step is significantly faster.

Production Gotchas & Troubleshooting

  1. Cache Misses Despite No Code Changes:

    • Symptom: Turborepo reports cache MISS even when no relevant code has changed.
    • Cause:
      • Incorrect inputs or outputs in turbo.json. If a file that influences the build is not listed in inputs, its changes won't invalidate the cache. If outputs are not correctly defined, Turborepo might not store the full output.
      • Environment variables not consistently set. If env variables defined in turbo.json change between runs, the hash changes.
      • package.json or pnpm-lock.yaml changes in a dependency.
      • fetch-depth in actions/checkout is too shallow, preventing Turborepo from accurately detecting changes.
    • Fix:
      • Review turbo.json inputs and outputs carefully. Use git diff --name-only <commit-ish> to identify files that changed and ensure they are covered.
      • Ensure all relevant environment variables are consistently passed to Turborepo.
      • Set fetch-depth: 0 in actions/checkout.
      • Run pnpm turbo run <task> --dry-run=json to inspect the computed hash and inputs.
  2. turbo prune Not Including Necessary Files:

    • Symptom: Docker build fails because a file or package expected by the target application is missing after turbo prune.
    • Cause:
      • The turbo prune command's --scope is too narrow, or the dependency graph is not correctly understood by Turborepo.
      • The COPY commands in the Dockerfile are incomplete, not copying all necessary source files from the original monorepo context.
    • Fix:
      • Verify the dependency graph using pnpm turbo graph --filter=api.... Ensure all required packages are listed.
      • Manually inspect the .pruned-build directory after running turbo prune.
      • Double-check the COPY commands in your Dockerfile, ensuring all source directories for the target app and its direct dependencies are included. Remember turbo prune only handles package.json and lockfiles, not source.
  3. GCS Permissions Errors:

    • Symptom: Turborepo fails to upload or download from GCS with permission denied errors.
    • Cause: The service account key used does not have sufficient IAM permissions on the GCS bucket.
    • Fix: Grant Storage Object Admin or at least Storage Object Creator and Storage Object Viewer roles to the service account on the specific GCS bucket. Ensure the TURBO_REMOTE_CACHE_GCS_SERVICE_ACCOUNT_KEY path is correct and the file is readable.
  4. Slow pnpm install in Docker:

    • Symptom: Even with turbo prune, pnpm install takes a long time in the Docker build.
    • Cause:
      • The pnpm-lock.yaml or package.json files are frequently changing, invalidating the Docker layer cache for pnpm install.
      • No pnpm cache is used in the Dockerfile.
    • Fix:
      • Ensure pnpm-lock.yaml is committed and kept up-to-date.
      • Structure your Dockerfile to copy pnpm-lock.yaml, package.json, and turbo.json first, then run pnpm install. This maximizes layer caching.
      • Consider using a shared pnpm cache volume for local development, though this is harder in CI. The actions/setup-node cache helps in CI.

Frequently Asked Questions

Q1: How does Turborepo handle changes in transitive dependencies for caching?

Turborepo's hashing mechanism considers the package.json of the current package and all its transitive dependencies. If any package.json in the dependency chain changes, or if the pnpm-lock.yaml (or yarn.lock, package-lock.json) changes, the hash for tasks depending on those packages will be invalidated, leading to a cache miss. This ensures correctness.

Q2: Can I use a different remote cache backend than GCS?

Yes. Turborepo supports various remote cache providers. Besides GCS, it natively supports Vercel Remote Cache (for Vercel deployments), AWS S3, and a generic HTTP endpoint. For S3, you would use TURBO_REMOTE_CACHE_S3_BUCKET, TURBO_REMOTE_CACHE_S3_REGION, and AWS credentials. For a custom HTTP endpoint, you'd configure TURBO_REMOTE_CACHE_URL.

Q3: What's the impact of fetch-depth: 0 on CI performance and security?

fetch-depth: 0 instructs Git to fetch the entire history of the repository, not just the latest commit.

  • Performance: For very large repositories with long histories, this can add a few seconds to the checkout time. However, it's often negligible compared to the time saved by effective Turborepo caching.
  • Security: Fetching the entire history means more data is pulled into the CI runner. For public repositories, this is generally not a concern. For private repositories, ensure your CI environment is secure. It's a necessary trade-off for Turborepo's accurate change detection.

Q4: How do I debug Turborepo cache issues locally?

Use the --dry-run flag with pnpm turbo run.

  • pnpm turbo run build --filter=api... --dry-run: Shows which tasks would run and why (e.g., cache MISS, cache HIT).
  • pnpm turbo run build --filter=api... --dry-run=json: Provides a detailed JSON output, including the computed hash for each task, its inputs, and outputs. This is invaluable for understanding why a cache miss occurred.
  • Set TURBO_LOG_LEVEL=debug for verbose logging.

Q5: My Docker image is still large despite turbo prune and multi-stage builds. What could be wrong?

  • Unnecessary files copied: Double-check your .dockerignore file. Ensure it excludes development-only files, .git, node_modules (except for the pnpm install step), and other build artifacts.
  • Build-time dependencies in runtime image: Ensure your production stage (runner in the example) only copies built artifacts and production dependencies. If you copy node_modules from the builder stage directly without re-installing with --prod, you'll include dev dependencies.
  • Large base image: node:20-alpine is generally lean. Avoid larger base images like node:20 (Debian-based) for production.
  • Leftover build tools: Ensure your runner stage doesn't contain compilers, linters, or other build tools from the builder stage. The example Dockerfile correctly addresses this by only copying specific dist directories and re-installing pnpm for prod dependencies.

By meticulously applying these strategies, organizations can achieve significant reductions in CI build times, leading to faster feedback loops, more frequent deployments, and improved developer productivity.

Share this article:

Stay Updated

Get the latest posts delivered straight to your inbox.

Free Developer Utilities

Free In-Browser Developer Tools

Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.

Explore Tools
Advertisement