Keep pulling the thread on Amjad Masad.
Replit's Agent 3 can autonomously test the software it writes by spinning up a browser, navigating the application, and iterating on the code to fix any issues it finds.
The duration for which Replit's AI agents can maintain coherence has increased from 2 minutes for Agent 1, to 20 minutes for Agent 2 (released in February), to 200 minutes for Agent 3.
To extend agent coherence to 200-300 minutes, Replit implemented a verifier in the loop where one agent can spin up a browser to perform automated testing on another agent's work.
Performance on the SWE-bench benchmark for software engineering has improved from approximately 5% in early 2024 to a state-of-the-art score of 82%.
Amjad Masad predicts that by next year, Replit users will be able to run multiple parallel agents simultaneously to perform tasks like planning new features and refactoring a database.
Amjad Masad predicts that very soon, a layperson using AI tools will be as productive as a senior software engineer at Google is today.
Ilya Sutskever has argued that AI progress may be limited by a shortage of training data, as most of the data on the internet has already been used.
Amjad Masad is bearish on a breakthrough to "true AGI" because current AI systems are so economically valuable that the industry is likely caught in a "local maximum" trap, reducing pressure to solve the general problem.
Replit's AI agents can take a paragraph-long English description of a startup idea and begin building it.
Replit's AI can automatically classify a user's request and select the best technology stack, such as Python and Streamlit for a data app or JavaScript and Postgres for a web app.
Amjad Masad concluded last year that Replit's business was underperforming because the requirement for users to write code with specific syntax was the primary bottleneck.
Replit experienced degraded performance for users in Asia after launching AI agents because the agents, acting as the "programmer," were located in the United States, creating high latency to servers in Asia.