Recent posts
[VLM] CommerceVibe: 이커머스 크리에이티브를 실행 가능한 HTML/CSS 코드로 생성하는 Dual-Feedback RL
[VLM] CommerceVibe: Learning to Design E-Commerce Creatives as Executable Visual Code via Dual-Feedback Reinforcement Learning
[LLM] On-Policy Self-Distillation: 하나의 LLM이 스스로의 teacher가 되어 추론을 가르치기
[LLM] On-Policy Self-Distillation for Large Language Models
[LLM] Magistral: 증류 없이 순수 RL만으로 만든 Mistral의 첫 추론 모델
[LLM] Magistral: 증류 없이 순수 RL만으로 만든 Mistral의 첫 추론 모델