warc-gpt

WARC + AI - Experimental Retrieval Augmented Generation Pipeline for Web Archive Collections.

Fork A fork of harvard-lil/warc-gpt; the README may describe the upstream project.

Overview

WARC + AI: Experimental Retrieval Augmented Generation Pipeline for Web Archive Collections.

https://github.com/harvard-lil/warc-gpt/assets/625889/8ea3da4a-62a1-4ffa-a510-ef3e35699237

WARC-GPT requires the following machine-level dependencies to be installed.

Use the following commands to clone the project and instal its dependencies:

This program uses environment variables to handle settings. Copy .env.example into a new .env file and edit it as needed.

See details for individual settings in .env.example.

Place the WARC files you would to explore with WARC-GPT under ./warc and run the following command to:

Note: Running ingest clears the ./chromadb folder.

The following command will start WARC-GPT's server on port 5000.

Once the server is started, the application's web UI should be available on http://localhost:5000.

Returns a list of available models as JSON.

Returns RAW text stream as output.

From the project’s README on GitHub.

At a glance

RepositoryKentucky-Open-Science/warc-gpt
Research areaForks of other projects
Primary languagePython
LanguagesPython 45.3%, JavaScript 33.2%, CSS 12.4%, Shell 6.2%, HTML 2.8%
LicenseMIT
Stars / forks0 / 0
Open issues and pull requests1
Created2024-07-24
Last push2025-10-22
Default branchmain
Homepagehttps://lil.law.harvard.edu/blog/2024/02/12/warc-gpt-an-open-source-tool-for-exploring-web-archives-with-ai/
Forked fromharvard-lil/warc-gpt — WARC + AI - Experimental Retrieval Augmented Generation Pipeline for Web Archive Collections.

What the README covers

Top contributors

Get the code

git clone https://github.com/Kentucky-Open-Science/warc-gpt.git