Repository navigation
Commit c7f4ba6
Fix out-of-order sample loss by making remote write synchronous
The store() method previously fired HTTP writes asynchronously via
executeAsync() and returned immediately. When the RingBuffer's
multiple worker threads dispatched consecutive batches containing
samples for the same series, the async HTTP requests could arrive
at the remote write endpoint out of timestamp order, causing the
backend to reject the stale samples as out-of-order.
This change makes store() block until the HTTP write completes,
ensuring the ring buffer worker thread does not process the next
batch until the current write has landed. This preserves per-series
timestamp ordering across consecutive WriteRequests as required by
the Prometheus Remote Write spec.
Additionally fixes a bug where samplesLost incorrectly counted
unfiltered samples (including NaN) instead of the actual samples
that were attempted.
Validated via A/B E2E testing against Thanos Receive:
- Baseline (async): 14 out-of-order / 5,262 appended (0.27% loss)
- Fix (sync): 0 out-of-order / 5,264 appended (0.00% loss)
- Throughput: identical (~5,260 samples over equal soak periods)
- All 45 smoke tests passing on both runs
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>1 parent fe9bda4 commit c7f4ba6
1 file changed
Lines changed: 13 additions & 9 deletions
Lines changed: 13 additions & 9 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
233 | 233 | | |
234 | 234 | | |
235 | 235 | | |
236 | | - | |
237 | | - | |
238 | | - | |
239 | | - | |
240 | | - | |
241 | | - | |
242 | | - | |
243 | | - | |
244 | | - | |
| 236 | + | |
| 237 | + | |
| 238 | + | |
| 239 | + | |
| 240 | + | |
| 241 | + | |
| 242 | + | |
| 243 | + | |
| 244 | + | |
| 245 | + | |
| 246 | + | |
| 247 | + | |
| 248 | + | |
245 | 249 | | |
246 | 250 | | |
247 | 251 | | |
| |||
0 commit comments