Started Updated
KTUGPT
Ask a question, get an answer from your KTU textbooks, with the page it came from.
- Python
- Flask
- LangChain
- Mistral 7B
- MongoDB Atlas
- Next.js
- Tailwind CSS
- Clerk

A final-year project
KTUGPT was our final-year B.Tech project in the Department of Information Technology at Government Engineering College Sreekrishnapuram. Our team was Muhammed Fariz KP, Nuzaim Noushad Thappi, Sameemul Haque K (me) and Sreelakshmi Jayarajan, guided by Dr. Rani M. R. We went through a few other ideas before we landed on this one.
The problem was one every KTU student knows. Exams follow the prescribed textbooks, but finding an answer in them means searching through the pages by hand. Google and ChatGPT answer quickly, but their answers are pulled together from many sources. They’re general answers, not the one your textbook gives.
So we set out to build an app that students and teachers could use to get answers directly from the KTU prescribed textbooks, with a reference to the page each answer came from.
Research first
We spent the first half of the project, from August to November 2023, on research: a literature survey, then setting our objectives, the design and the work plan.
The papers we read shaped the approach. One of the choices was between fine-tuning a model on the textbooks and in-context learning. Fine-tuning needs labelled training data. In-context learning gives a pre-trained model the relevant text as part of the prompt and asks it to answer from that. We had raw textbooks, not labelled data, so we went with in-context learning combined with embeddings. Today this approach is usually called retrieval-augmented generation (RAG).
How it works

- Preparing the books. The textbook PDFs are split into chunks of text. Each chunk is turned into an embedding (a list of numbers that captures its meaning) and stored in MongoDB Atlas, which we used as our vector database.
- Finding the right pages. When someone asks a question, the question is turned into an embedding too. A similarity search finds the three chunks closest in meaning.
- Answering. Those chunks and the question go to the language model, with a prompt that tells it to answer only from that context and to say it doesn’t know if the answer isn’t there.
- Showing the source. The answer comes back with the textbook name and page number for each chunk, so you can open the book and read more.
Picking the models
We tried several models and compared them on hardware we could actually afford.
For the language model, we chose Mistral 7B Instruct. It responded faster than the others, and it stuck to the context we gave it:
- Llama 7B Chat was slower and needed around 30 GB of RAM.
- Falcon 7B Instruct sometimes gave answers that weren’t really about the given context.
- distilbart-cnn-12-6 is much smaller, at 306 million parameters, and its results were unpredictable.
For embeddings, we chose hkunlp/instructor-xl, which takes task-specific instructions into account when it builds an embedding:
- OpenAI’s text-embedding-ada-002 was good but costs money.
- all-MiniLM-L6-v2 can only take 256 words of input.
- all-mpnet-base-v2 was slower.
Building it
The backend is a Flask service that uses LangChain to connect the embeddings, MongoDB Atlas and the language model. It was deployed on a Hugging Face Space. The frontend is built with Next.js and Tailwind CSS, uses Clerk for email and OTP login, keeps a chat history for each user, and shows the references under every answer.

We split the work like this:
- Data collection and preparation: Fariz, Sameemul and Sreelakshmi.
- Frontend: Fariz and Sreelakshmi.
- Backend: Nuzaim and Sameemul.
- Everyone together: choosing and configuring the models, testing, and the documentation.
My main part was the backend: the Flask service, the retrieval and prompt, returning the source pages, and deploying it. On the frontend side, I also added the Clerk login, worked on the chat history, connected the UI to the backend and made it responsive.
What we found
We compared KTUGPT’s answers with Google’s and ChatGPT’s for the same questions. Theirs were fine, but general. KTUGPT’s came straight from the textbook, and every answer pointed to the book and page, so it could be trusted and checked.
It had clear limits too:
- It could only generate text, so diagrams and numerical problems were out.
- It worked for direct answers, not analytical questions that need deeper understanding.
- It was computationally expensive. Because of our hardware limits, we only loaded a few chapters from two KTU prescribed Operating Systems textbooks: Silberschatz’s Operating System Concepts and Tanenbaum’s Modern Operating Systems.
With the right cloud storage and a GPU, anyone could load more textbooks and run it for other subjects.
It was my first real project with large language models and embeddings. Looking back, it’s where I learned how RAG works, well before it became something every app seems to have.