APIs for document processing and retrieval.
A user-scoped FastAPI backend for document intelligence. It processes PDF, DOCX, TXT, CSV and XLSX files, stores OpenAI embeddings in PostgreSQL with pgvector, and generates summaries and retrieval-grounded answers with source citations. Spreadsheet analytics, downloadable PDF reports and durable SSE processing events support the document workflow.
01 / SYSTEM MAP
[ CONNECTED LAYERS ]- 1APIAuthenticated document servicesFastAPI / JWT / Pydantic
- 2RETRIEVALEmbeddings & cited answersOpenAI / PostgreSQL / pgvector
- 3PROCESSINGAnalysis, reports & progresspandas / ReportLab / SSE
02 / MY CONTRIBUTION
Backend architecture & AI pipeline development
- +Implemented JWT authentication and owner-scoped document, folder and tag APIs.
- +Built upload validation, background text extraction, chunking and vector storage for supported document formats.
- +Connected semantic retrieval to AI summaries and cited Q&A with follow-up question history.
- +Added spreadsheet statistics, cached AI insights and downloadable PDF reports.
- +Implemented database-backed status events and authenticated SSE streams for processing updates.
subdirectory_arrow_rightPublished source, not a live API. The current implementation uses FastAPI background tasks and local upload/report storage.