Use large language model to enhance reasoning of another large language model through reward updated GRPO.
Journal:
Scientific reports
Published Date:
Feb 11, 2026
Abstract
No abstract available for this article.
Authors
Keywords
No keywords available for this article.