I want to interrogate a database of blog posts using AI but ChatGPT has token limits

stinky@redlemmy.com · 1 year ago

I want to interrogate a database of blog posts using AI but ChatGPT has token limits

hedgehog@ttrpg.network · 1 year ago

Retrieval-Augmented Generation (RAG) is probably the tech you’d want. It basically involves a knowledge library being built from the documents you upload, which is then indexed when you ask questions.

NotebookLM by Google is an off the shelf tool that is specialized in this, but you can upload documents to ChatGPT, Copilot, Claude, etc., and get the same benefit.

If you self hosted, Open WebUI with Ollama supports this, but far from the only one.

theunknownmuncher@lemmy.world · 1 year ago

Dunno why this is downvoted because RAG is the correct answer. Fine tuning/training is not the tool for this job. RAG is.

Danitos@reddthat.com · 1 year ago

OP can also use an embedding model and work with vectorial databases for the RAG.

I use Milvus (vector DB engine; open source, can be self hosted) and OpenAI’s text-embedding-small-3 for the embedding (extreeeemely cheap). There’s also some very good open weights embed modelsln HuggingFace.

Scrubbles@poptalk.scrubbles.tech · 1 year ago

I understand conceptually how these work, but I have a hard time of how to get started . I have the model, I know embeddings exist and what they are, and rags, and vector dbs, and then I have my SQL DB. I just don’t know what the steps are.

Do you have any guides you recommend?

Danitos@reddthat.com · 1 year ago

Milvus documentation has a nice example: link. After this, you just need to use a persistent Milvus DB, instead of the ephimeral one in the documentation.

Let me know if you have further questions.

Scrubbles@poptalk.scrubbles.tech · 1 year ago

That’s a great start! A lot of it depends on OpenAI, is there any guide you know of that lets me run completely locally? I use TabbyAPI for most of my inference, and happy to run anything else for training

Danitos@reddthat.com · edit-2 1 year ago

It would work the same way, you would just need to connect with your local model. For example, change the code to find the embeddings with your local model, and store that in Milvus. After that, do the inference calling your local model.

I’ve not used inference with local API, can’t help with that, but for embeddings, I used this model and it worked quite fast, plus was a top2 model in Hugging Face. Leaderboard. Model.

I didn’t do any training, just simple embed+interference.

Scrubbles@poptalk.scrubbles.tech · 1 year ago

Ah okay, I think that makes sense. Thanks for your input! I’ll give it a whirl

stinky@redlemmy.com · 1 year ago

Thanks, I got NotebookLM working pretty quickly. I think RAG is what I’m after. I’ll continue to look.

manicdave@feddit.uk · edit-2 1 year ago

If you want to try the openwebui route, This guide might be helpful.

Edit: in fact I don’t think this is for openwebui specifically, but I remember the chapter at the timestamp is what helped me increase the context window. That’s the important bit if you’re wanting to ask it questions about documents.