Author
MIT researchers propose Supervised Memory Training, which uses predictive-state labels from a Transformer encoder to train nonlinear RNN updates in parallel without backpropagation through time.
We use cookies to improve your experience and analyze site traffic. You can choose which cookies to allow. Privacy Policy