Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Contrastive Decoding Diffing (CDD): recovering verbatim finetuning data from logits alone, no weight access needed[R]

Via r/MachineLearning
Friday, Jul 3, 2026 · 7:01PM
Summary

We built a model diffing method that recovers verbatim content from narrowly finetuned LLMs using only grey-box logit access (no weights, no activations, no probe corpus). Recent work (Minder, Dumas et al., "Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences") showed that fin

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories