Amanda Askell, mentioned 43 times across podcast episodes and expert conversations analyzed by Sonic.
▶Both Askell and Anthropic CEO Dario Amodei agree that the personality and character of an AI model are deliberately engineered. Askell describes the goal for Claude's personality as being nuanced and charitable, originating from an alignment initiative, while Amodei discusses character training to manage user expectations and the risks of personality changes in updates.Jul 2026
▶There is a shared understanding that direct instructions are a key tool for shaping model behavior. Askell details the use of system prompts as a rapid, low-cost patching method for specific issues, which aligns with Amodei's discussion of using instructions to make the model explain its own limitations.Jul 2026
▶Both acknowledge that post-training is crucial for shaping the final model. Askell specifies that reinforcement learning primarily elicits latent abilities from the pre-trained model, a view compatible with Amodei's focus on the overall goal of the training process, which is to build robust and safe systems.Jul 2026
▶Both highlight the importance of managing specific, undesirable model behaviors. Askell details efforts to handle sycophancy and political bias, while Amodei discusses the ongoing work to reduce overly apologetic behavior in Claude.Jul 2026
▶Askell's claims focus on the specific mechanics of AI training, such as the role of system prompts, the function of RLAIF, and the iterative nature of prompt engineering. In contrast, claims from her colleague Dario Amodei in the same conversation address broader, strategic concerns like defining AGI, the philosophy of building robust systems, and the societal risks of user-AI attachment.
▶Askell emphasizes that post-training techniques primarily 'elicit' latent capabilities already present in the pre-trained model. Amodei's focus is less on this mechanistic distinction and more on the end-goal of the entire process: ensuring the final model is safe enough to prevent catastrophic outcomes.Jul 2026
▶Askell provides granular detail on specific behavioral fixes implemented via system prompts, such as ensuring symmetrical treatment of political topics or preventing the model from claiming objectivity. Amodei discusses the higher-level user experience trade-offs involved in such changes, for example, noting that making a model less apologetic increases the risk of it appearing rude when it makes an error.Jul 2026
Create a free account to see Amanda Askell's full intelligence report - every claim, the relationship network, and AI Q&A across all sources. No card needed.
Get started free