Self-Hosting

Troubleshooting

Resolve common boot, run, network and configuration problems when self-hosting Appstrate.

Start with docker compose logs appstrate --tail 200 (or appstrate logs for an installer install) and curl http://localhost:3000/health. Boot errors name the variable at fault.

A stack written by appstrate install runs under a derived Compose project name. For it, use appstrate logs and appstrate status, or add --project-name "$(jq -r .projectName .appstrate/project.json)" to the docker compose commands on this page.

The container will not start

Five secrets are required and validated at boot:

VariableRule
BETTER_AUTH_SECRETRequired
CONNECTION_ENCRYPTION_KEY32 bytes, base64-encoded
UPLOAD_SIGNING_SECRETAt least 16 characters
RUN_TOKEN_SECRETAt least 16 characters
CONNECT_SESSION_SECRETAt least 16 characters

The shipped Compose files also fail fast on POSTGRES_PASSWORD, on MINIO_ROOT_PASSWORD in the files that run MinIO (the root file, the Tier 3 template and examples/self-hosting/docker-compose.yml), and on POSTGRES_USER and MINIO_ROOT_USER in the root file. See Docker Compose for a generator.

Other cross-field checks that stop the boot:

  • TRUST_PROXY=false behind an HTTPS APP_URL in production. Set TRUST_PROXY to your proxy depth, usually 1.
  • APP_URL is not an absolute origin. It must be http:// or https:// with no path, query or fragment, and HTTPS in production (loopback excepted).
  • USERCONTENT_URL on the same host as APP_URL, or not HTTPS in production.
  • Image versions disagree. The platform, PI_IMAGE and SIDECAR_IMAGE must carry the same version. Pin them with APPSTRATE_VERSION. See Upgrading.
  • PLATFORM_RUN_LIMITS, INLINE_RUN_LIMITS, LLM_PROXY_LIMITS or CREDENTIAL_PROXY_LIMITS contains an unknown key or a bad value. The error names the key.
  • RUN_ADAPTER names a backend nobody registered. The error lists the registered ones. firecracker needs the firecracker module in MODULES.
  • A model in SYSTEM_PROVIDER_KEYS is outside its provider's offer, or a provider needs an API shape the platform does not serve. The error names the entry.

A variable in .env has no effect

Two common causes:

  • Compose does not forward it. The container only receives the variables named under appstrate.environment in the Compose file. Add a bare - NAME line. Even the current files do not list every variable (LLM_PROXY_LIMITS and FILE_MAX_BYTES are two that they do not), and a file from before 1.0.0-beta.65 lacks more. See Compose forwards only what it lists.
  • The name is wrong. Appstrate does not recognise renamed variables, and an unknown key is stripped without a warning. Compare the spelling with Environment Variables and the operator notes of the release.

Database connection refused

  • Check that DATABASE_URL points to a reachable PostgreSQL, and that the database container is healthy: docker compose ps.
  • On the Docker tiers the database sits on an internal network and publishes no port. Test from inside: docker compose exec postgres pg_isready (service appstrate-postgres in the root file).
  • If DATABASE_URL is unset, Appstrate uses PGlite in PGLITE_DATA_DIR. A single PGlite directory must be opened by one process only.
  • A password with @, / or : breaks the connection string Compose builds. Use hex or URL-safe passwords.

Agents fail to run

  • Wrong backend. RUN_ADAPTER defaults to process (host subprocesses, no isolation). The shipped Compose files set docker.
  • Docker socket. The platform container must reach /var/run/docker.sock. Check the group mapping (DOCKER_GID) and the socket's permissions. See Isolation and Security.
  • Runtime images missing. Check docker images | grep appstrate. Pull the same version of appstrate-pi, appstrate-sidecar and the five appstrate-mcp-runner-* images from ghcr.io/appstrate. A missing MCP runner image fails the integration that needs it, because the sidecar never pulls images.
  • Local integration refused under process. A source.kind: "local" integration (for example @appstrate/github-git) does not start with RUN_ADAPTER=process. Set INTEGRATION_RUNTIME_ADAPTER=docker or use RUN_ADAPTER=docker.
  • "Run never started executing". The runner posted no event within RUN_BOOT_DEADLINE_SECONDS (default 300). Look for a slow image pull, a missing image or a container that exits at boot.
  • Run failed as stalled. No heartbeat for RUN_STALL_THRESHOLD_SECONDS (default 60). Check the run's container logs and the host's resources.
  • "Server restarted while run was in progress". A restart finalizes in-flight runs as failed. Retry them.
  • Model on a private endpoint fails. Add its host to EGRESS_ALLOW_INTERNAL_HOSTS. The host must resolve from the API process, where localhost is the API's own loopback.

Outbound calls refused

SymptomMeaning
403 URL targets a blocked network range, and Target refused (SSRF) in the sidecar logThe target is private, loopback or link-local. To allow it, name the host literally in the integration's authorized_uris and list it in EGRESS_ALLOW_INTERNAL_HOSTS
403 with unauthorized_target (platform credential proxy), or a sidecar refusal as unauthorizedThe integration declares no authorized_uris, or the URL is outside them. An empty list authorizes nothing
502 Target host could not be resolvedThe sidecar could not resolve the host. It needs working DNS, even when it sends through PROXY_URL

See Isolation and Security.

Docker network pool exhausted

Docker create network appstrate-exec-... failed: 400
{"message":"all predefined address pools have been fully subnetted"}

Each run uses one network, and Docker's default address pool holds about 31 on a stock host. Reclaim unused networks with docker network prune. For a lasting fix, carve smaller subnets in /etc/docker/daemon.json (Docker Desktop: Settings, Docker Engine) and restart the daemon:

{
  "default-address-pools": [
    { "base": "172.20.0.0/16", "size": 24 },
    { "base": "10.200.0.0/16", "size": 24 }
  ]
}

Each /16 base then yields about 256 networks. Appstrate also reclaims orphaned appstrate-exec-* networks after a failure and retries once.

400 on API calls: missing organization or space

Requests that belong to an organization need an X-Org-Id header, and space-scoped routes (agents, runs, schedules, end-users, API keys, notifications, packages, integrations, files and uploads) also need X-Space-Id. Without it you get 400. An API key carries its own organization and space, so it needs neither header. With an API key, X-Org-Id is ignored (the key is bound to its organization) and only X-Space-Id is checked: a value that contradicts the key's space returns 403.

Sign-in and first-owner problems

  • Sign-up is closed (signup_disabled). The instance runs in closed mode. Ask for an invitation, or see AUTH_MODES.md.
  • /claim answers 410. The bootstrap token is only redeemable while the instance has no organization. If you set AUTH_BOOTSTRAP_TOKEN and it does nothing, check that your Compose file forwards the variable to the container (the files shipped before 1.0.0-beta.65 do not).
  • The owner address cannot register. An address named in AUTH_BOOTSTRAP_OWNER_EMAIL or AUTH_PLATFORM_ADMIN_EMAILS is not created by the plain sign-up form. Claim it at /claim while the instance has no organization (with AUTH_BOOTSTRAP_OWNER_EMAIL set, the claim accepts that address only), or use a magic link (SMTP required) or a Google or GitHub sign-in whose provider asserts the address as verified. The server log states the reason.
  • A magic link or a Google or GitHub sign-in fails right after an upgrade. Better Auth 1.7.7, shipped in 1.0.0-beta.65, changed how these in-flight values are stored. A link mailed before the restart is refused and a social sign-in started before it must be started again. Ask for a new link. See Upgrading.
  • Redeem rate-limited. /api/auth/bootstrap/redeem allows 5 attempts per minute per IP. Behind a proxy, set TRUST_PROXY.

OAuth callback fails

APP_URL must match the public URL exactly, including the scheme and no stray port:

# correct
APP_URL=https://appstrate.example.com

# wrong: missing scheme
APP_URL=appstrate.example.com

# wrong: internal port when a proxy terminates TLS on 443
APP_URL=http://appstrate.example.com:3000

Register the same redirect URI at the provider (Google, GitHub or an OIDC client).

Uploads and live streams misbehave behind a proxy

  • 413 that Appstrate never logs. The proxy's body limit is lower than your upload. Raise it to at least 100 MiB (nginx: client_max_body_size 100m).
  • Chat or run logs arrive in one batch at the end. The proxy buffers or compresses text/event-stream. Disable that for the Appstrate location.
  • Everyone is rate-limited together. TRUST_PROXY is not set to your proxy depth, so every caller shares the proxy's IP. See Rate Limits.

Redis fallbacks

Without REDIS_URL (Tier 0 and 1), the queue, pub/sub, cache and rate limiter are in-process:

  • The scheduler's cron evaluator polls every 30 seconds, and queued jobs are lost on restart.
  • Rate-limit counters reset on restart.
  • Several instances do not share state. Set REDIS_URL for more than one instance.

With Redis, keep its volume if you care about scheduled runs and queued webhook deliveries.

MinIO crash-loops

FATAL Unable to initialize backend: Unable to write to the backend means the miniodata volume is not owned by uid 65532. Re-own it once with the stack stopped. The exact commands, with a snapshot step first, are in the self-hosting README.

Stored objects that will not delete

Deleting a file, a run workspace, a space or an organization removes its database rows at once and queues the stored objects in a deletion outbox. A background worker purges them. Every STORAGE_DELETION_WORKER_INTERVAL_MS (default 60000, see Environment Variables) it claims a batch of due jobs and deletes the objects. A failed job is retried later, with a delay that doubles after each attempt (about one minute at first, capped at 6 hours, with up to 10% jitter). A job is never abandoned: deletion is retried for as long as it takes.

After 8 attempts a job that is still pending is called a dead letter. The threshold only makes the job visible, it does not stop the retries. Dead letters are what to look at when disk or bucket usage does not go down after deletions. With the @appstrate/module-observability module, the gauges appstrate.storage_deletion.backlog, appstrate.storage_deletion.oldest_pending_age_seconds and appstrate.storage_deletion.dead_letters report the same state.

List the jobs with the platform-admin API:

curl "https://your-instance/api/admin/storage-deletion-jobs?status=dead" \
  -b cookies.txt
Query parameterValues
statuspending (default, every job not yet completed), dead (pending with 8 or more attempts) or completed
limit1 to 200, default 50
startingAfterThe id of the last job of the previous page. Follow the Link header with rel="next" instead of building it by hand

The answer is { "object": "list", "data": [...], "hasMore": false }, newest first. Each job has id, bucket, storage_key (the key inside the bucket), reason (why the object is purged, for example file_deleted, space_deleted or run_workspace_deleted), attempts, next_attempt_at, completed_at, last_error (the message of the last failed attempt) and createdAt. The list covers every organization of the instance, and a storage key contains a space id and a file name.

To retry a job without waiting for its next scheduled attempt:

curl -X POST "https://your-instance/api/admin/storage-deletion-jobs/$JOB_ID/retry" \
  -b cookies.txt

It answers { "id": "...", "retried": true } and makes the job due immediately, so the next worker pass picks it up. A job that is already completed, or an unknown id, answers 404. Retrying does not reset the attempt count, so a dead letter stays one until it succeeds. The two calls are limited to 60 and 30 per minute. Fix the cause shown in last_error first (the storage backend, its credentials or its permissions), or the retry fails the same way.

Both calls answer 403 Platform admin access required unless all of these hold:

  • You are signed in with a dashboard session cookie. An API key is refused, and so is an OIDC token, whatever its scopes.
  • The session belongs to the platform audience. A person who signed in as an end-user of a space is refused.
  • Your email address is in AUTH_PLATFORM_ADMIN_EMAILS, a comma-separated list compared without regard to case. With the variable empty, nobody qualifies. The account for a listed address is only created with proof that you own it, see AUTH_MODES.md.

These two routes need no X-Org-Id, so an operator who belongs to no organization can call them. Before 1.0.0-beta.65 they answered 400 without that header.

Rate limit hits (429)

Check Retry-After and the RateLimit headers. Per-endpoint limits and the organization-wide run limits are in Rate Limits. An organization at its concurrent run cap gets org_run_concurrency_exceeded.

Port 3000 is already in use

Change the host side of the port mapping, or set PORT in .env (the Compose files read ${PORT:-3000}). appstrate install --port 3100 sets it for you, and with --yes the installer picks the next free port by itself.

minisign: command not found during install

curl -fsSL https://get.appstrate.dev | bash verifies the CLI binary with minisign. On a TTY it offers to install minisign through your package manager, and with --yes, in CI or without a TTY it installs it automatically when it can. If that fails, install it yourself (brew install minisign, apt install minisign, apk add minisign) and re-run. Setting APPSTRATE_NO_INSTALL_MINISIGN=1 disables the automatic install.

Verbose logging

LOG_LEVEL=debug

Debug logs include one access line per request (method, route pattern, status, duration and Request-Id). Return to info once the incident is understood.

Orphaned containers

Containers created for runs carry the label appstrate.managed=true, and the platform reconciles them at startup. If strays remain after a crash:

docker ps -a --filter "label=appstrate.managed=true" -q | xargs docker rm -f

Per-run networks (appstrate-exec-*) are cleaned up at boot. The shared appstrate-egress network is never removed, by design.

On this page