Lil'Log
Summary
Lil'Log covering Reinforcement Learning from Human Feedback (RLHF), AI Safety, Reward Hacking, and Reinforcement Learning (RL). Notable guests include Lilian Weng. Episodes span from Jun 2023 to Jul 2026.
AI agent performance is increasingly dependent on 'harness engineering'—the system of workflows, tools, and memory management surrounding a base model—which is becoming as critical as the model's c...
Model performance in deep learning, particularly for Transformers, scales predictably as a power law with increases in compute, model size (N), and dataset size (D). A central debate, resolved by t...
Increasing 'test-time compute' or 'thinking time' through methods like Chain-of-Thought (CoT) is a critical technique for enhancing the reasoning capabilities of large language models, especially f...