How can you prove that a training set was used to train a neural network?
That's the question considered by Choi, Shavit and Duvenaud in the work "Tools for Verifying Neural Model's Training Data".
Timestamps:
00:00 - Tools for Verifying Neural Model's Training Data
00:24 - Why should we care about this problem?
02:42 - Formal Definition for Proof-of-Training Data
04:57 - Using checkpoints as a witness
06:18 - Uniqueness and Faithfulness
07:25 - Verification strategies
08:03 - Memorization-based tests
10:32 - Fixing the Initialization and Data Order
12:15 - Putting it all together
12:47 - Experimental Setup
13:02 - Gluing attack
13:39 - Interpolation attack
14:02 - Data addition attack
14:28 - Data subtraction attack
14:43 - Discussion and Limitations
15:41 - Broader Impacts
16:30 - Closing thoughts
Topics: #verification #auditing #proof-of-training-data
Link to the paper: https://arxiv.org/abs/2307.00682
For related content:
Twitter: / samuelalbanie
Research lab: https://caml-lab.com/
personal webpage: https://samuelalbanie.com/
YouTube: / @samuelalbanie1
TikTok: / samuelalbanie
Instagram: / samuelalbanie
LinkedIn: / samuel-albanie
Threads: https://www.threads.net/@samuelalbanie
Discord server for filtir: / discord
(Optional) if you'd like to support the channel:
https://www.buymeacoffee.com/samuelal...
/ samuel_albanie