---
title: Limits
description: Quotas, fleet bounds, rate limits, and the scale-to-zero sleep mechanism.
---

## Quotas

| Knob | Default | Effect |
| --- | --- | --- |
| `RUNNER_MAX_APPS_PER_USER` | `10` | Per-account app quota, enforced in the runner on create — API keys cannot bypass it |
| `RUNNER_BUILD_MAX_MB` | `2048` | Builds refused past this worktree size (MiB) |
| `RUNNER_BUILD_TIMEOUT_S` | `300` | Each install and build step is killed past this (seconds) |
| `RUNNER_BUILD_CACHE_MB` | `512` | Per-app build cache cap (Bun, npm, pnpm, Yarn and their jup-managed releases); over it, whole caches are dropped, largest first; `0` disables |
| `RUNNER_TELEMETRY_RETENTION_DAYS` | `30` | Telemetry lifetime in bucket prune, metric tables, and glob floor |

## Fleet bounds

Every app is its own celld fleet on one host, sharing one cgroup. The bounds stop one account — or one busy app — from owning the machine.

:::warning
`RUNNER_FLEET_MAX_RSS_MB` is container-wide, not per-app: celld compares it against the greater of its own RSS and the cgroup working set, and every fleet shares the one container cgroup. A per-fleet-sized value closes every fleet's admission gate as soon as the container as a whole passes it. Size it to the container; leave `0` for celld's default (80% of container memory).
:::

| Knob | Default | Effect |
| --- | --- | --- |
| `RUNNER_FLEET_MAX_RSS_MB` | `0` | Shed threshold in MiB; over it, fleets refuse cells with `503` + `Retry-After` instead of OOM-ing |
| `RUNNER_FLEET_IDLE_EVICT_S` | `120` | Idle seconds before a fleet hibernates cells and returns memory |
| `RUNNER_FLEET_ASSET_CACHE_MB` | `64` | Per-fleet on-disk asset cache (celld default is 512) |
| `RUNNER_FLEET_DEPLOY_POLL_S` | `300` | Fleet self-poll for deploys; the runner also POSTs `/reload`, so this covers misses only |
| `RUNNER_FLEET_LOG` | `error,celld=warn` | Fleet log filter; celld warnings reach the per-app log view |

## Rate limits

The edge (Caddy) limits every site before a request reaches a worker; the control UI adds finer limits of its own on sign-in routes. Requests per minute, `0` disables any of them:

| Layer | Var | Default | Scope | Over-limit |
| --- | --- | --- | --- | --- |
| Edge, per visitor | `NOITE_EDGE_RPM` | `1800` | Per client IP, per site: control, API, the fallback page, each app across all its hostnames | `429` + `Retry-After` |
| Edge, git | `NOITE_EDGE_GIT_RPM` | `120` | Per client IP on `git.` | `429` + `Retry-After` |
| Edge, whole app | `NOITE_EDGE_APP_RPM` | `0` (none) | Per app, across every client | `429` + `Retry-After` |
| Worker limiter | `NOITE_RATE_LIMIT_RPM` | `600` | Per client, per route class (`auth`/`invite`) | `429` + `Retry-After`, before better-auth or D1 |
| better-auth budget | `NOITE_AUTH_RATE_LIMIT` | `600` | Per client on `/api/auth/*`, keyed on forwarded address | `429` |

An app admin can override both app limits (per visitor and whole app) in the app's **Settings → Rate Limits**. The edge counts over a sliding 10-second window: a limit of 1800/min lets 300 requests through in any 10 seconds. Per-visitor limits group IPv6 addresses by `/64`. Behind a proxy, see [Edge protection](/self-hosting/protection).

## Sleep (scale to zero)

An app with no requests for `RUNNER_SLEEP_AFTER_H` hours (default 24, `0` disables) stops costing anything; the next request brings it back transparently.

**Sweep.** Every `RUNNER_SLEEP_SWEEP_S` seconds (default 3600) the runner parks apps that are deployed, desired `running`, awake, and request-free for the window.

**Activity** is the newest of: the last minute bucket with requests in `app_metric` (real `celld.fetch` spans), the app's last wake (`woke_at`), its last deploy. A fresh deploy or manual start gets a full window.

**Asleep is not stopped.** `desired_state` stays `running`; sleep is its own column `asleep_since`, and status reads `sleeping`. A stopped app never sleeps or wakes; stopping an asleep app clears the flag. Going to sleep: mark asleep → rewrite Caddy sites to wake-on-demand → stop the fleet with normal SIGTERM drain.

**Wake flow.** Asleep sites route to the fleet port preceded by `forward_auth` to `GET /v1/edge/wake`. Caddy holds the request, body included, while the runner:

1. Resolves the app from `X-Forwarded-Host` (tenant slug or custom domain).
2. Clears the flag, sets `woke_at`.
3. Spawns the fleet immediately.
4. Waits for `/.well-known/celld/health` 200, bounded by `RUNNER_WAKE_TIMEOUT_S` (default 120).
5. Rewrites the Caddyfile back to the plain proxy.
6. Answers 200; Caddy proxies the held request as if the app never slept.

Concurrent requests share one wake (per-app lock), and at most `RUNNER_WAKE_CONCURRENCY` apps (default 4) cold-start at once; the rest queue within the same timeout. Only a failed wake answers `503`. A deploy, rollback, or web commit to an asleep app wakes it first; TLS ask treats asleep apps as live so certificates renew. First-request cost after a quiet day: one cold fleet start (celld boot + ready gate, typically seconds).
