Projects

Things I built and measured

Two systems, each with the result that came out of it — including the one that contradicted what I set out to show.


Selected projects

Retrieval Evaluation Benchmark

Read the write-up

A single-command evaluation harness — modular indexing, fusion, reranking, scoring — comparing BM25, dense, hybrid RRF, and cross-encoder retrieval over three BEIR datasets (34K docs, 1,623 judged queries), with 71 offline tests.

Cross-encoder reranking degraded nDCG@10 on 2 of 3 datasets (−0.034, −0.023; CIs exclude zero) at 70× the query latency, and hybrid fusion lost to dense alone where the lexical input was weak.

Implemented rank-based Reciprocal Rank Fusion, a dependency-free Porter stemmer matching the Anserini BM25 analyzer, and paired bootstrap 95% confidence intervals over 1,000 resamples. Failure analysis disproved the standard exact-match hypothesis, isolating query–document lexical overlap as the real mechanism; hand-written nDCG cross-validated against ir_measures to exact agreement over 1,300 queries.

Python · Sentence-Transformers · BEIR · ir_measures · pytest · 2026

Agentic multi-omics platform

omicsai.org ↗

A React frontend backed by FastAPI and a Gradio AI interface over an agentic multi-omics system, with LangGraph_Orchestrator coordinating multi-agent workflows across the platform.

Shipped to the lab's demo platform on AWS EC2, serving papers under review at Oxford Bioinformatics.

React · FastAPI · Streamlit · Gradio · LangGraph · Docker Compose · AWS EC2 · Caddy (HTTPS, subpath routing)