Person
1 entry · 23 December 2020
The system learned to predict only value, policy and reward rather than the environment's rules, then matched AlphaZero at Go, chess and shogi and set a new Atari benchmark.
Models & capabilities