{"author_name":"Paper Radio","author_url":"https://m.montaan.com/","height":720,"html":"<iframe src=\"https://m.montaan.com/embed/video/df3b5a90-3949-4a7b-a4c3-e1766de0a24f\" width=\"1280\" height=\"720\" title=\"AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning — Paper Radio\" frameborder=\"0\" allow=\"autoplay; fullscreen; picture-in-picture\" allowfullscreen loading=\"lazy\" style=\"border:0;display:block;max-width:100%;\"></iframe>","provider_name":"Paper Radio","provider_url":"https://m.montaan.com/","thumbnail_url":"https://m.montaan.com/media/episodes/df3b5a90-3949-4a7b-a4c3-e1766de0a24f/episode.poster.webp","title":"AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning","type":"video","version":"1.0","width":1280}