Keep pulling the thread on Eamonn Maguire.
Anthropic is reportedly facing a $1.5 billion lawsuit for allegedly scanning thousands of purchased books for training its Claude models and then destroying the physical copies.
Meta was found to have used a large archive of pirated books to train its AI models.
Eamonn Maguire cites the case of Molly Russell, who committed suicide at age 14, as an example where Instagram's algorithm allegedly propagated harmful content related to suicide.
Proton was founded at CERN in 2014, inspired by the Edward Snowden revelations.
Proton was initially crowdfunded through a Kickstarter project that raised approximately $500,000.
Proton is funded by its subscribers and does not have any venture capital investors.
Proton operates on a freemium model, where paying subscribers fund the free services available to non-paying users.
Free users of large language model platforms like ChatGPT and Claude typically give permission by default for their data to be used for model training.
Enterprise users of LLM platforms typically have contracts stating their data will not be used for model training, but the speaker believes a trust issue remains as the companies still have access to the data.
NVIDIA is creating its NeMo models based on open data.
The entire English-language section of Wikipedia constituted only 0.3% of the training data for OpenAI's GPT-2 model.
Eamonn Maguire believes it is likely that OpenAI is seeking access to the research material in the Bodleian Library at Oxford University for model training.