Replication lag shows up as stale reads on followers or slow recovery after a node restart. Diagnose before you restart observers.
Symptoms
replica lagalerts from OCP or custom queries.- Clients see inconsistent reads on follower routing.
- Backup jobs skip tablets citing lag.
Diagnosis
SELECT ls_id, svr_ip, role, replay_lag_ms, append_lag_ms
FROM oceanbase.__all_virtual_log_stat
ORDER BY replay_lag_ms DESC
LIMIT 20;
Check network between lagging follower and leader. Inspect disk util on follower; 90%+ iowait stalls replay.
Mitigation
- Throttle bulk load on leader tenant.
- Temporarily remove follower from read pool.
- If log disk full, expand
log_disk_sizeand restart observer (planned window).
Escalate if lag > 300s and growing. This may indicate Paxos partition.