Person
1 entry · October 2023
Sparse autoencoders decomposed a single 512-neuron layer into more than 4,000 human-interpretable features, far more than the raw neurons showed.
Safety & alignment