Video wird geladen...
Video konnte nicht geladen werden
Anthropic published a way to read LLM’s inner thoughts last week (Natural Language Autoencoders). 2 days later, we spent 36h straight hours at Platanus Ventures Hack Buenos Aires, were we built an open-source system that uses it to detect deception and steer models back into alignment. Here’s what we found:
15,943 Aufrufe • vor 4 Monaten •via X (Twitter)
0 Kommentare
Keine Kommentare verfügbar
Kommentare vom Original-Post werden hier angezeigt
