ALIGNMENT
What an Honest Agent Looks Like, and Why We Have So Few
Honesty in language models was, until recently, treated as a matter of training data. The new generation of work suggests it is closer to a matter of architecture.
Read article →