native_research_llms

Registry of LLMs pretrained from scratch by universities, non-profit research institutes, and national laboratories

Overview

An index of foundation models trained natively, from scratch, on their own data and tokenizer, by academic and research institutions. Universities come first; research institutes, national labs, and public consortia are tracked alongside.

Foundation model development is dominated by well-funded commercial labs, and data opacity, licensing constraints, and architectural gatekeeping come with it. In response, universities and other public institutions have pretrained large language models from a tabula rasa state: random weights, own corpus, own tokenizer, auditable end to end.

Entries marked ⚠ carry a caveat worth reading before you rely on them, usually a restrictive license or a documented access constraint.

No commercial partner, no corporate-donated compute. University teams on university or public compute. This is the headline ranking.

From the project’s README on GitHub.

At a glance

RepositoryKentucky-Open-Science/native_research_llms
Research areaLanguage models & AI tooling
Primary languageRuby
LicenseCC0-1.0
Stars / forks3 / 0
Open issues and pull requests0
Created2026-07-16
Last push2026-09-22
Default branchmain
Topicslarge-language-models

What the README covers

Top contributors

Get the code

git clone https://github.com/Kentucky-Open-Science/native_research_llms.git