Keep pulling the thread on Dan Biderman and Jessy Lin.
Engram's continual learning approach can reduce token inference consumption by up to two orders of magnitude (100x) by internalizing context that would otherwise require large system prompts.
A model trained with Engram's technology can answer certain queries in 100 tokens, whereas a frontier model might consume 100,000 tokens for the same task.
Engram aims to use offline computation to compress large KV caches, potentially reducing an 80-gigabyte cache by a factor of 1,000x.
Dan Biderman's vision is for Engram to become the primary LLM interface to the data plane for all users, drawing a parallel to the roles of companies like Databricks and Oracle.
Engram is working with partners including Notion, Microsoft, and Harvey to develop models that deeply understand context within their workspaces.
Engram's technical approach involves training per-team models within partner workspaces that utilize adapter fine-tuning techniques like LoRAs, prefixes, and sparse architectures.
Engram's technology requires white-box access to model weights, making it most easily applicable to open-source models.
Engram considers individual computers and phones to be future targets for its continual learning technology.
Dan Biderman argues that if an AI lab like OpenAI needed to win a math Olympiad within a week, the superior strategy would be to synthesize training data and launch a new training job rather than relying on a retrieval-based system.
Engram's vision for the future of AI involves many personalized models for individuals and teams, which contrasts with the Frontier Lab worldview of building a single, increasingly large general model.
Demis Hassabis stated at a Sequoia event approximately one month prior to this interview that new breakthroughs are needed around the topics of AI memory and continual learning.
Some of the best large language models from China incorporate layers inspired by state-space architectures, which allows them to operate with sub-quadratic cost.