Intelligent Chatbot Service / Document AI & RAG
Answers grounded in your documents.
Document ingestion, retrieval, and source-aware answers behind configurable chatbots.
Document parsing preserves layouts, tables, and attribution as content enters the retrieval pipeline.
Illustration of the document-to-answer pipeline, not a screenshot of the private product.
- LangChain
- ltree
- AWS Lambda
- Amazon S3
- DynamoDB
- Amazon SQS
- RAGAS
- Tonic Validate
- OpenTelemetry
- AWS X-Ray
Results & scope
- Average document-processing time per page
- 0.5 s
- CV project result
- Reported document-parser result in the December 2024–August 2025 chatbot project.
Method & conditions
Self-reported in my CV. Hardware and document mix are not specified. No comparison to another parser is implied.
- Reported processing cost per page
- $0.0007
- CV project result
- Cost reported for the document-processing pipeline.
Method & conditions
Self-reported in my CV. The cost accounting boundary is not specified. Not a current hosting-price estimate.
- Pages processed continuously
- 400–500
- CV project result
- Reported continuous document-processing runs without failure.
Method & conditions
Self-reported in my CV. Run scale for the recorded project, not a guaranteed limit for every document type.
- Reported cost per 1,000 SEC records
- $0.26
- CV project result
- SEC-filing extraction service in the December 2024–August 2025 chatbot project, combining a self-hosted document parser with LLM extraction.
Method & conditions
Self-reported in my CV. The cost accounting boundary and input mix are not specified. This is a historical project result, not a current service quote or a verified comparison with commercial APIs.
- Reported FinQA answer similarity
- 90%
- CV project result
- Document-question answering evaluation reported in my CV, alongside RAGAS and Tonic Validate evaluation work.
Method & conditions
Self-reported answer similarity, not answer accuracy. The CV does not specify the dataset version, evaluated split, sample count, or similarity calculation. No benchmark-leadership claim is made.
- Reported FinanceBench answer similarity
- 70%
- CV project result
- Document-question answering evaluation reported in my CV, alongside RAGAS and Tonic Validate evaluation work.
Method & conditions
Self-reported answer similarity, not answer accuracy. The CV does not specify the dataset version, evaluated split, sample count, or similarity calculation. No benchmark-leadership claim is made.
How I approached the work
Engineering decisions.
Keep document structure and sources through retrieval
- Problem
- Useful document answers depend on more than raw text: layouts, tables, and source attribution must survive ingestion.
- Decision
- Built structure-aware processing behind serverless queues and customized LangChain ingestion, retrieval, and generation. Used ltree hierarchies with vector similarity and weighted scoring for document-based recommendations.
- Result
- Configurable document chatbots with contextual retrieval, source-aware processing, and evaluation using RAGAS and Tonic Validate.
The work.
I developed a platform where users upload documents and configure a chatbot around their information. My work connected document parsing and event-driven ingestion with retrieval, generation, evaluation, and distributed tracing.
My contribution
Documents into context
Built ingestion and parsing for document layouts, tables, and source attribution, including SEC-filing extraction, backed by serverless processing queues.
Retrieval with control
Customized LangChain ingestion and retrieval. Combined ltree hierarchies with vector similarity and weighted scoring to recommend insights from extracted document data.
Evaluation and tracing
Integrated RAGAS and Tonic Validate evaluation, plus OpenTelemetry and AWS X-Ray to inspect behavior across services.