Aryan Kargwal is a passionate developer, researcher, and advocate in the field of AI, specializing in large language models (LLMs), vision-language models (VLMs), and generative AI. His work bridges the gap between cutting-edge technology and practical applications, empowering developers and organizations to harness the power of AI effectively. He is currently a PhD student at PolyMTL.
The latest from Aryan Kargwal


TLDR: An agent, by definition, can take actions and is more than just a language model. The model is only one part of the system, and is only one element that you are training. The environment determines what the agent can observe, what it can change, which actions are available, and what behavior receives a reward. For a coding agent,…
Learn how RLVR uses verifiable rewards to train models, how training data and verifier design shape performance, and where RLVR still struggles with agents and long-horizon tasks. TL;DR Reinforcement learning with verifiable rewards (RLVR) has emerged as a practical approach for training reasoning models on tasks with objectively checkable outcomes. RLVR can provide consistent training feedback without requiring a human…



