Use large language model to enhance reasoning of another large language model through reward updated GRPO.

Journal: Scientific reports
Published Date:

Abstract

No abstract available for this article.

Authors

Keywords

No keywords available for this article.