Person
Jacob Hilton
Appears alongside
Featured in threads
Tracks
- Safety & alignment 2
- Ideas & essays 1
- Labs & people 1
- Benchmarks & progress 1
Current and former staff demand a right to warn
Thirteen current and former employees of OpenAI, Google DeepMind and Anthropic signed; Bengio, Hinton and Russell endorsed it without being employees themselves.
Ideas & essays · Safety & alignment · Labs & people
TruthfulQA measures whether models repeat human falsehoods
On 817 questions designed to elicit common misconceptions, the best model tested was truthful only 58% of the time against 94% for humans, and larger models scored worse.
Benchmarks & progress · Safety & alignment