Charge-payload cleanup
This runbook is for the append-only Rivian charging payload evidence table. The
normal synchronizer should reuse a semantic payload identity and should not
rewrite unchanged sessions or curve points. Cleanup is bounded and dry-run by
default; it is not a substitute for fixing a write-amplification regression.
The canonical charge_payload_fingerprint and
charge_payload_identity_key database functions are shared by ingestion,
backfill, and this compactor; do not reintroduce an inline identity expression.
Retention-aware identities
rivian_charge_payloads is TimescaleDB evidence retained for 90 days. Its
JSON rows are intentionally dropped by the normal retention policy; this is
not a database corruption event. The separate payload identity table keeps the
semantic identity key and fingerprint indefinitely, but its
canonical_payload_id is only an optional pointer to retained evidence.
When the last matching JSON payload has expired, the identity becomes a
logical tombstone (canonical_payload_id = NULL). Replaying the same upstream
session must not insert replacement JSON or repeatedly grow the hypertable.
The session reconciliation path still runs normally without a payload pointer.
The background identity worker runs a small cache-reference reconciliation pass
after its historical backfill completes and then every six hours. It clears
orphaned identity pointers and all orphaned alias cache pointers independently;
it never deletes identities, changes the 90-day policy, or manufactures payload JSON.
An aggregated charge_payload_retention_reconciled INFO log reports the
identity and alias counts. A poll-level aggregated INFO with
payloads_evidence_expired means retained evidence expired; investigate it
only if it affects unexpectedly recent sessions. A poll-level aggregated WARN
indicates a repair or payload_audit_failures.
Before cleanup
- Create and verify a recovery package or PostgreSQL dump.
- Record the current payload, alias, session, and curve-point counts.
- Pause the Riviamigo app/worker or otherwise schedule the operation away from active charging-history synchronization. The compactor takes its own advisory lock to prevent concurrent compactors, but the safest window is still a quiet ingestion window.
- Confirm that the target database is the intended instance and that the backup has been copied off the NAS.
Diagnose first
Run without --apply:
cargo run --manifest-path apps/api/Cargo.toml --bin compact_charge_payloads -- \
--batch-size 5000
The report includes relation and payload-byte totals plus the number of
semantic duplicates. It does not delete rows. Use --vehicle <uuid> to bound
the report to one vehicle.
Continuous PostgreSQL activity should be investigated before compacting. Check
the app logs for charge history synced counters, including payloads_reused
and payloads_inserted, and compare the counts across two idle polls. Normal
checkpoint, vacuum, or Redis persistence activity is not evidence of a
charging-history loop by itself.
Apply a bounded cleanup
After reviewing the dry run, remove one bounded batch:
cargo run --manifest-path apps/api/Cargo.toml --bin compact_charge_payloads -- \
--batch-size 5000 --apply
The compactor keeps a linked/oldest canonical payload, repoints aliases and semantic identities, and deletes only redundant payload rows. Repeat the same bounded command while the report shows remaining candidates. Stop if the reported identity, alias, or session counts change unexpectedly.
Do not use VACUUM FULL on a live production database. After the bounded
deletion and a quiet window, run a normal maintenance vacuum/analyze according
to your PostgreSQL operations policy. VACUUM FULL, pg_repack, or a dump and
restore require additional outage space and rollback planning; select one only
when the recovered disk space justifies that operational cost.
After cleanup
- Re-run the dry-run report and record zero remaining semantic duplicates (or the reviewed remainder for a filtered vehicle).
- Compare payload, alias, session, and curve-point counts with the pre-cleanup record.
- Restart the stack and verify
/health, worker status, and a signed-in dashboard. - During the next idle completed-history poll, confirm
payloads_reusedincreases whilepayloads_insertedremains zero for identical upstream data. For sessions older than 90 days,payloads_evidence_expiredmay increase instead; it must not cause new payload rows. - For an active charge, confirm that curve writes correspond only to new or genuinely corrected points, not the full historical curve length.
If counts or costs change unexpectedly, stop synchronization, restore the verified backup in an isolated database, and investigate the replay before deleting more evidence.
Rollback compatibility
Migration 0014_charge_payload_retention_references only makes
canonical_payload_id nullable and documents the cache semantics. It performs
no data rewrite and does not alter TimescaleDB retention. A rollback to code
that assumes a non-null canonical UUID is unsafe once tombstones exist; first
deploy a version that understands nullable pointers, or restore the compatible
database backup in an isolated environment.