AI systems and large language models need to be trained on massive amounts of data to be accurate but they shouldn’t train on data that they don’t have the rights to use. Human Native AI is a London-based startup building a marketplace to broker such deals between the many companies building LLM projects and those willing to license data to them. It’s goal is to help AI companies find data to train their models on while ensuring the rights holders opt in, where will all the training data be sourced from ?

