Scaling LyftData
Use this playbook when jobs start to backlog, worker utilisation spikes, or new workloads arrive. Start from evidence: adding capacity changes only what is eligible to run where; it does not establish safe input ownership, recovery, or failover.
Know when to scale
Section titled “Know when to scale”Monitor these signals (see Monitoring LyftData for how to collect them). The numbers below are example investigation triggers, not universal limits. Replace them with thresholds derived from a measured healthy period for each workload.
| Metric | Action threshold | Typical response |
|---|---|---|
| Worker CPU usage | > 80% sustained | Profile the job, then add eligible capacity or CPU |
| Worker memory usage | > 85% sustained | Check leaks/state, then adjust memory or batch size |
| Job queue length | > 100 pending | Check input/downstream health, ownership, and scheduling |
| Job runtime | 2× baseline | Profile actions, add capacity |
| Error/retry rate | > 5% | Investigate downstream systems before scaling |
Prove that horizontal scaling is safe
Section titled “Prove that horizontal scaling is safe”Before increasing replicas or adding workers, answer these questions:
- How does the input assign partitions or records to consumers?
- Is state local, shared, or reconstructable, and where are checkpoints stored?
- Must records remain ordered or deduplicated?
- Can the chosen transport cross workers and recover after interruption?
- Can every eligible worker reach the source, destination, credentials, and storage?
- What runtime evidence proves improvement without loss or duplication?
If any answer is unknown, measure and design that boundary before calling the change safe. See Placement and Scaling, Transports and Durability, and Delivery Semantics.
Horizontal capacity (add workers)
Section titled “Horizontal capacity (add workers)”- Provision new worker hosts or containers as close to your data sources as possible to minimise latency.
- For licensed deployments, create a worker ID in Workers and an API key in Settings → API Keys (see Workers).
- Point the worker at the control plane and supply credentials:
Terminal window lyftdata-worker \--url https://lyftdata.example.com \--worker-id worker-02 \--worker-api-key "$LYFTDATA_WORKER_API_KEY" - For fleets that rely on auto enrollment, configure auto-enrollment on the server, then start each worker with
LYFTDATA_AUTO_ENROLLMENT_KEYinstead of pre-issuing API keys. - Verify the worker registers successfully via the UI or
lyftdata workers list.
Worker configuration checklist
Section titled “Worker configuration checklist”LYFTDATA_URL– HTTPS URL for the control plane.- Identity – either
LYFTDATA_WORKER_ID+LYFTDATA_WORKER_API_KEY, orLYFTDATA_AUTO_ENROLLMENT_KEYwith an optionalLYFTDATA_WORKER_NAME. LYFTDATA_WORKER_LABELS– optional comma-separated descriptors surfaced in the UI for filtering and documentation.--limit-job-capabilities– start the worker with this flag when it should only accept basic connectors (handy for hardened ingress zones).- Systemd or your orchestrator should manage the worker process so it restarts automatically after upgrades or crashes.
Workers advertise their capabilities back to the server. The UI shows each worker’s advertised limits (including maximum concurrent jobs). In Community Edition, scaling remains constrained by built-in-worker-only operation, licensed-feature gating, and the daily processing cap.
Vertical scaling (bigger boxes)
Section titled “Vertical scaling (bigger boxes)”When horizontal scale is not practical:
- Increase CPU/memory allocations for the server or existing workers.
- Ensure disks that host the staging directory have headroom; adjust
--db-retention-daysor--disk-usage-max-percentageif cleanup is too aggressive. - Revisit batch sizes and action efficiency so additional resources translate into throughput.
Automate worker scaling
Section titled “Automate worker scaling”- Poll server health and worker status on a schedule (via the UI, CLI, or API) and persist readings so you can compare moving averages instead of reacting to single spikes.
- Feed those metrics into your autoscaling tool (Kubernetes HPA, AWS ASG, Nomad autoscaler, etc.). A typical policy adds a worker when queue depth stays above your threshold for several minutes.
- During scale-in, first drain or reassign input ownership according to the connector or broker contract. Stop workers gradually and watch for backlog, duplicates, ordering changes, and errors.
- Alert whenever automation fails (for example, API calls start returning errors or workers flap repeatedly) so operators can intervene manually.
Placement strategies
Section titled “Placement strategies”- Region-aware: co-locate workers with data sources to reduce egress and latency.
- Capability-aware: dedicate workers to specialised connectors or sensitive networks and label them so runbooks show where to deploy specific jobs.
- Spare capacity: keep eligible capacity in another failure domain only after proving the source, state, storage, credentials, and placement can recover there. Spare capacity alone is not automatic failover.
Keep jobs efficient
Section titled “Keep jobs efficient”Scaling is easier when jobs are lean:
- Split large multi-purpose jobs into smaller stages with Workflows. Use worker channels only when both jobs are deliberately co-located; use a durable transport when stages cross workers or need replay.
- Use
filterearly to drop unnecessary records. - Profile Lua scripts and external lookups; cache results where possible.
- Adjust scheduling cadence (see Advanced scheduling) to smooth bursty workloads.
Review after scaling
Section titled “Review after scaling”- Validate that job runtimes and queue lengths return to normal.
- Compare input, output, duplicate, and error counts with the pre-change baseline; use Runtime Evidence and Receipts and destination readback for external effects.
- Update documentation/runbooks with the new topology.
- Revisit resource budgets monthly; scale down when demand shrinks to save cost.
Pair this guide with the backup and security runbooks to keep operations predictable as the platform grows.