Recent posts
[Agent] MENTOR: Teacher-Optimized Reward로 Tool-Use를 SLM에 distill하는 on-policy RL
[Agent] MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation
Agent Distillation
title: “[Agent] Agent Distillation: CoT가 아니라 ‘행동’을 증류해서 0.5B 모델을 도구 쓰는 에이전트로 만들기”
[VLM] CommerceVibe: 이커머스 크리에이티브를 실행 가능한 HTML/CSS 코드로 생성하는 Dual-Feedback RL
[VLM] CommerceVibe: Learning to Design E-Commerce Creatives as Executable Visual Code via Dual-Feedback Reinforcement Learning