DeepMind Releases ReST, an Algorithm to Improve Translation Quality

Google DeepMind released a paper proposing ReST, an algorithm that makes it simpler to align LLM with human preferences.Unlike RLHF (Reinforcement Learning Based on Human Feedback) which uses human feedback to improve language models, ReST is trained by generating and using offline data.

Previous:

Next:

Leave a Reply

Please Login to Comment