-
-
Notifications
You must be signed in to change notification settings - Fork 1
Runbooks
github-actions[bot] edited this page Sep 6, 2026
·
1 revision
Troubleshooting guides for common failure scenarios. Full runbooks in docs/runbooks/.
Symptoms: Pipeline status = FAILED, events not flowing
Diagnosis:
# Check pipeline status
curl -H "Authorization: Bearer $TOKEN" http://localhost:8080/api/pipelines/{id}
# Check DLQ
curl -H "Authorization: Bearer $TOKEN" http://localhost:8080/api/dashboard/errors
# Check logs
kubectl logs -l app=syncflow --tail=100 | grep -i "pipeline.*error"Resolution:
- Verify source/destination connections are healthy
- Check for schema changes in source database
- Review DLQ for failed events
- Restart capture:
POST /api/pipelines/{id}/capture/stopthenstart
Symptoms: CDC lag > 30 seconds, events delayed
Diagnosis:
# Check CDC metrics
curl http://localhost:8080/api/diagnostics/connectors
# Check capture status
curl -H "Authorization: Bearer $TOKEN" http://localhost:8080/api/pipelines/{id}/capture/statusResolution:
- Increase
syncflow.runtime.sync.batch-size - Check destination write latency
- Verify network between source and SyncFlow
- Consider adding Kafka transport for cross-pod distribution
Symptoms: Snapshot resumes from wrong position, duplicate data
Diagnosis:
# Check snapshot checkpoints
psql -d syncflow -c "SELECT * FROM snapshot_checkpoints WHERE pipeline_id = '{id}'"Resolution:
- Delete corrupted checkpoints:
DELETE FROM snapshot_checkpoints WHERE pipeline_id = '{id}' - Restart snapshot from beginning
- Verify
chunk_indexvalues are sequential
Symptoms: Agent not sending heartbeats, jobs failing
Diagnosis:
# Check agent status
curl -H "Authorization: Bearer $TOKEN" http://localhost:8080/api/agents
# Check agent logs
kubectl logs -l app=syncflow-agent --tail=100Resolution:
- Verify agent can reach control plane URL
- Check agent token is valid
- Restart agent pod
- Drain agent before shutdown:
POST /api/agents/{id}/drain
Symptoms: OOMKilled, heap exhaustion
Diagnosis:
# Check JVM metrics
curl http://localhost:8080/api/diagnostics/system
# Check heap usage
curl http://localhost:9090/actuator/metrics/jvm.memory.usedResolution:
- Increase
-Xmxin JVM args - Reduce
syncflow.runtime.sync.queue-capacity - Reduce
syncflow.runtime.snapshot.parallelism - Check for memory leaks in connector clones
Symptoms: Write failures, Flyway migration errors
Diagnosis:
# Check database size
psql -d syncflow -c "SELECT pg_size_pretty(pg_database_size('syncflow'))"
# Check table sizes
psql -d syncflow -c "SELECT schemaname, tablename, pg_size_pretty(pg_total_relation_size(schemaname||'.'||tablename)) FROM pg_tables WHERE schemaname='public' ORDER BY pg_total_relation_size(schemaname||'.'||tablename) DESC"Resolution:
- Clean DLQ:
DELETE FROM dead_letter_events WHERE created_at < NOW() - INTERVAL '7 days' - Clean processed events:
DELETE FROM processed_events WHERE created_at < NOW() - INTERVAL '30 days' - Vacuum:
VACUUM FULL ANALYZE - Add storage or archival policy
Symptoms: Sync latency > 60 seconds
Diagnosis:
# Check sync metrics
curl http://localhost:8080/api/dashboard/metrics
# Check writer performance
curl http://localhost:8080/api/diagnostics/connectorsResolution:
- Increase
syncflow.runtime.sync.batch-size - Check destination database performance
- Verify network latency
- Consider connection pooling (already uses HikariCP)
- Check for lock contention in destination
Symptoms: Events retrying frequently, approaching DLQ threshold
Diagnosis:
# Check retry metrics
curl http://localhost:8080/api/dashboard/errors
# Check specific pipeline errors
curl -H "Authorization: Bearer $TOKEN" http://localhost:8080/api/pipelines/{id}/capture/statusResolution:
- Identify root cause from error messages
- Fix transient issues (network, locks)
- Adjust
syncflow.runtime.retry.max-attemptsif needed - Review DLQ for permanent failures
Symptoms: DLQ growing, events not processing
Diagnosis:
# Check DLQ count
curl http://localhost:8080/api/dashboard/errors
# List DLQ events
curl -H "Authorization: Bearer $TOKEN" http://localhost:8080/api/dead-letter?pipelineId={id}Resolution:
- Fix root cause of failures
- Replay resolved events:
POST /api/dead-letter/{id}/replay - Purge old events:
DELETE /api/dead-letter?olderThan=7d - Set up alerts for DLQ depth threshold