What are you looking for ?
VergeIO
RAIDON

Corbenic AI Launches Galahad: the GPU Never Reads the Same Text Twice Again

A memory layer for AI that runs inside vLLM, SGLang and llama.cpp. Free on one GPU

When a model reads a text, the GPU does expensive work to understand it. Today, that work is thrown away and done again every time the same text comes back: the same contract, the same manual, the same chat history. Galahad keeps that work. When the text is needed again, Galahad hands the computed information back instead of letting the GPU read it a second time. Galahad is also built to find the parts a question is about, so the model only reads what matters.Galahad plugs easily into the three most used open-source inference engines for running proprietary AI models: vLLM, SGLang and llama.cpp. Galahad saves up to 98.7% of the GPU’s computational work. In tests on seven real-world, messy datasets, it retrieved 98.7% of the text from its memory using 97% less energy per answer.

“A GPU should never pay twice for the same reading. Galahad remembers what your model has already read and hands it back. Everything we claim is measured. With the free version anyone can check it on their own GPU,” said Sietse Schelpe, founder, Corbenic AI.

Availability
Free for non-commercial use on one GPU, for 12 months and renewable.

Articles_bottom
AIC