Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

๐Ÿ‘ผ BASICS OF ETHICAL HACKING โ€“ CORE TRAINING ๐Ÿ˜ˆ ๐Ÿ”ฐ Learn key topics like CEH overview, hacking concepts, security threats, attack types, and building a safe hacking lab. ๐Ÿ”— LINK:

12,232 Aufrufe โ€ข vor 6 Monaten โ€ขvia X (Twitter)

0 Kommentare

Keine Kommentare verfรผgbar

Kommentare vom Original-Post werden hier angezeigt

ร„hnliche Videos

Arena intern and UCLA PhD candidate, Hengguang Zhou, introduces Trace-and-Amplify (TA), a framework for collecting training-time reward-hacking trajectories at scale without hacking instructions. Monitors trained and evaluated on prompt-elicited hacking trajectories can achieve high detection accuracy, but often fail to transfer to training-time reward-hacking trajectories that emerge during RL without hacking instructions. Trace-and-Amplify enables scalable collection of these training-time trajectories, producing monitors that generalize much better to real and held-out hacking types. Detection accuracy 59.98% (PE-trained) โ†’ 90.16% (TA-trained) compared to 97.1% on prompted hacks โ†’ 28.0% on training-time hacks. 0:00 โ€“ OpenAI's ExploitGym cyberattack benchmark exploit 1:04 โ€“ Goodhart's Law and the CoastRunners boat-racing hack (2016) 2:04 โ€“ Gaming the evaluator: the robot-hand grasping example (2017) 3:10 โ€“ Reward hacking in code generation: hard-coding, test-rewriting, skipping eval 4:20 โ€“ A standard defense: reward-hacking monitors 4:58 โ€“ Monitor architectures: zero-shot LLMs, fine-tuned BERT, hidden-state probes 6:11 โ€“ Where monitor training data comes from today: prompted hacks 7:03 โ€“ The core question: do prompted hacks represent real hacks? 7:35 โ€“ Why this matters: RL post-training is the standard recipe for frontier models 8:20 โ€“ Why nobody's checked this before (hacking is rare, labeling isn't scalable) 9:23 โ€“ Introducing the method: Trace-and-Amplify 9:49 โ€“ The Tracer: a contradictory unit test that locates evaluation-gaming 10:50 โ€“ Amplify: collecting hacking rollouts at scale during RL training 11:32 โ€“ Experiment setup: Qwen2.5-Coder, DeepSeek-Coder, LeetCode/TACO 12:16 โ€“ Finding #1: prompt-trained monitors don't transfer to real hacks 14:40 โ€“ Can strong zero-shot judges (GPT-4.1, o4-mini) do better? 15:48 โ€“ Finding #2: monitors trained on real hacks generalize much better to unseen hacks 17:04 โ€“ Ruling out artifacts introduced by the method 17:55 โ€“ Why the gap? Real hacking is more hidden than prompted hacking 20:05 โ€“ Three takeaways, limitations, and future work

Arena.ai

33,585 Aufrufe โ€ข vor 1 Monat