Replication lag shows up as stale reads on followers or slow recovery after a node restart. Diagnose before you restart observers.

Symptoms

  • replica lag alerts from OCP or custom queries.
  • Clients see inconsistent reads on follower routing.
  • Backup jobs skip tablets citing lag.

Diagnosis

SELECT ls_id, svr_ip, role, replay_lag_ms, append_lag_ms
FROM oceanbase.__all_virtual_log_stat
ORDER BY replay_lag_ms DESC
LIMIT 20;

Check network between lagging follower and leader. Inspect disk util on follower; 90%+ iowait stalls replay.

Mitigation

  1. Throttle bulk load on leader tenant.
  2. Temporarily remove follower from read pool.
  3. If log disk full, expand log_disk_size and restart observer (planned window).

Escalate if lag > 300s and growing. This may indicate Paxos partition.