[Agent] MENTOR: Teacher-Optimized Reward로 Tool-Use를 SLM에 distill하는 on-policy RL
[Agent] MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation
[Agent] MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation
title: “[Agent] Agent Distillation: CoT가 아니라 ‘행동’을 증류해서 0.5B 모델을 도구 쓰는 에이전트로 만들기”
[VLM] CommerceVibe: Learning to Design E-Commerce Creatives as Executable Visual Code via Dual-Feedback Reinforcement Learning
[LLM] On-Policy Self-Distillation for Large Language Models
[LLM] Magistral: 증류 없이 순수 RL만으로 만든 Mistral의 첫 추론 모델