An internal chatbot for attachment-based reasoning and web search
Operating project
Document and image analysis and web research share one conversation interface. A request queue controls inference on limited GPU resources.
What the user sees
Simulation
Public demo answer detail · Not real inference
①InputQuestion · attachment selection
②AnswerStreaming display · feedback
③HistoryConversations · reopen
What the server handles
01Question & filesReact
02Auth & contextFastAPI
03Request queuePostgreSQL FIFO
04Model inferenceSGLang · GLM-5.3 Flash
05Streaming answerSSE → React
API ↔ PostgreSQL: accounts · conversations · files · feedback
A summary of the operating source. The public demo does not execute this server path.
02 / PERMISSION-AWARE RAG
RAG with a centralized internal reference database as the source for answers
Design baseline
Approved organizational documents become searchable knowledge. The design retrieves evidence within user permissions and links answers to their sources.
03Grounded answerDecline if evidence is insufficient
04Source referenceDocument · version · as-of date
Design concept: index approved documents, retrieve permitted evidence and attach sources to answers. Detailed ERD and implementation status are separated in the atlas.
Fictional format example
Question
Which documents are needed for expense reimbursement?
Answer
An application and supporting receipts are required. [1]