Streaming DataInf
Streaming DataInf is the proposal to use influence estimates to identify which recent examples affect a target behavior.
Project status: Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.
Mechanism
Approximate training-example influence in the low-rank adapter setting and maintain useful statistics as feedback arrives. The intended application is debugging harmful updates and selecting useful replay examples.
Implementation and controls
The roadmap and dedicated streaming-LoRA survey propose an influence probe before production integration. It needs a defined target loss, adapter parameter set, damping/curvature approximation, refresh cadence, memory budget, and a counterfactual test that removing a high-ranked event changes the target as predicted.
Evidence and evaluation
The cited project document records this candidate and its intended experiment. It does not provide a completed local result for this method. Published-paper results mentioned by that document are background, not Trainfer measurements.
Limitations and interpretation
An influence score is an approximation, not a causal guarantee. Streaming nonstationarity, optimizer history, and merges can invalidate cached quantities. No active DataInf API or controlled local influence-validation result was identified in the audited code.
Sources
- cont: docs/research/production-implementation-roadmap.md — checkout audited
87946914c7b9. - cont: docs/research/surveys/datainf-streaming-lora.md — checkout audited
87946914c7b9.