Keep pulling the thread on Mark Chen.
AI models have become significantly better at discovering novel theorems and advancing the frontiers of hard sciences.
The narrative that 'pre-training is dead' is incorrect as scaling laws continue to hold.
Scaling laws for LLMs have held for almost 10 orders of magnitude and are expected to continue holding with more careful research, data engineering, and scaling.
The field of AI is in an 'eval crisis' because canonical gold-standard benchmarks, such as the SAT, are now fully saturated by models.
A major organizational goal for OpenAI's research division is to develop models capable of self-sustained research and innovation.
Unlocking the primitive of continual learning is necessary for AGI, and it is highly likely that one of the many techniques being explored will be successful.
OpenAI's 3-year research roadmap aims to develop models capable of performing end-to-end research, which includes solving the problem of imbuing them with 'good taste'.
OpenAI has hired many researchers without formal training in machine learning or AI research, believing in training them internally.
The best mechanism for developing research taste in AI is to fully replicate admired academic papers.
AI models are now capable of performing long-horizon, meaningful work in professions like math, computer science, and coding.
Reinforcement Learning (RL) has traditionally struggled in subjective fields like creative writing where expert opinions can vary widely and grading is difficult.
Reinforcement Learning (RL) is most effective in fields with objective, hard truths like math and computer science, where correctness is clearly defined.