Production AI Systems: With the Receipts
Three deployable AI systems that solve real problems, each shipped with its evaluation harness, its failure taxonomy and the bugs that were found, including bugs in the evaluator itself. Helio is a cited, schema-validated solar advisory RAG that abstains instead of guessing. AssetPilot is a bounded multi-tool agent with hard budgets, five stop conditions and a per-step audit trail. Clearline is a data pipeline that validates messy multi-source exports into an AI-ready store with a per-run quality score.
79.2%RAG accuracy, zero false answers
10 / 10Agent tasks incl. 3 failure injections
343xError in a naive load the pipeline catches
PythonPydanticBM25LLM (live + mock)TerraformCI evals18 test gates
paretoeval
Cost-aware evaluation for LLM applications. Scores quality per dollar, attributes cost across every pipeline stage including retries, models self-hosted GPU economics with utilisation, and fails the build when cost or quality regresses. Composes with Ragas, DeepEval, tokencost and OpenTelemetry backends.
Pythonpytest pluginOpenTelemetryApache 2.0
Powerline: Alexa+ home power copilot
A stateful MCP server that makes a household solar inverter conversational. Answers "can I run the iron for 30 minutes?" with a real forecast, plans the day's appliances around solar surplus, learns the local grid pattern and pushes server-initiated low-battery alerts. Supports Growatt and Victron hardware with a built-in simulator. Built for the Amazon Developer Hackathon 2026.
TypeScriptMCP (Streamable HTTP)MCP AppsAmazon Bedrock38 tests
Blind-Zone TVT Prediction in Horizontal Wells
Kaggle ROGII Wellbore Geology Prediction. A change of coordinates turns a noisy regression into a tracking problem; a second-order hidden-state tracker follows structural position and dip through a banded trellis, fused with drift-cancelling and spatial-field branches. Runs on one CPU core at about four seconds per well. Entered with one month left in the competition.
241of 6,125 teams
+3,039places at private reveal
45%error reduction vs baseline
PythonNumPyBeam searchSignal processing
BirdCLEF+ 2026: Late-Entry Solution and Ensemble-Diversity Study
Multi-taxon acoustic species identification in the Brazilian Pantanal. Entered solo with two weeks left on free-tier compute. Extracted Perch v2 embeddings over 10,658 soundscapes, trained a custom head with pseudo-labelling, and documented a controlled negative result: apparent ensemble diversity on a shared frozen representation lives in noise, not complementary signal. The code behind an accepted CLEF 2026 working note.
0.9406private ROC-AUC
1,313 / 4,092final rank
+119places at private reveal
PythonPerch v2BioacousticsPseudo-labelling
Amadioha: Open-Domain Question Answering for Civic Engagement
A two-stage question answering system over a curated corpus of 1.3 million Nigerian news articles: BM25 retrieval in Apache Solr followed by a fine-tuned BERT span extractor. Includes linguistic error analysis on Nigerian English, bias evaluation of answer availability across entities, and an explainability audit so every answer shows its source. Presented at the Black in AI Workshop, NeurIPS 2019.
BERTApache SolrBM25NLP
Cognitive Visual Reasoning Agent for Raven's Progressive Matrices
A hybrid symbolic-perceptual agent that solves visual IQ puzzles from raw images. Contour extraction and object segmentation feed a relational representation; a shape-agnostic transformation engine covering rotation, reflection, XOR and progression drives a generate-and-test loop modelled on analogical reasoning and rule-based inference.
PythonOpenCVKnowledge-Based AI
Comparative Study of Supervised Learning Algorithms
Decision trees, random forests, gradient boosting, SVMs, neural networks and kNN benchmarked across datasets with hyperparameter sweeps, learning curves, bias-variance analysis and cross-validation, with attention to interpretability versus performance.
scikit-learnPython
Reinforcement Learning and Markov Decision Processes
Policy iteration, value iteration and Q-learning implemented from scratch and evaluated across reward structures, transition dynamics, discount factors and exploration schedules, with visualization of convergence and policy behaviour.
PythonMDPsQ-learning
Unsupervised Learning and Dimensionality Reduction
k-means, Gaussian mixture models, ICA, PCA, t-SNE and random projection applied across two datasets, with reduced representations fed into downstream neural network training to test when each technique recovers meaningful latent structure.
scikit-learnPython
Randomized Optimization
Genetic algorithms, simulated annealing and randomized hill climbing compared across optimization landscapes, then applied to neural network weight tuning, with analysis of temperature schedules and exploration versus exploitation.
PythonMetaheuristics
AI Fairness and Bias Auditing with IBM AIF360
Audited demographic bias in two real-world datasets, Taiwan credit allocation and the Mental Health in Tech survey, using statistical parity difference, disparate impact and consistency metrics, then applied pre-processing mitigation and measured the accuracy-fairness trade-off.
AI Fairness 360Python
Word2Vec and Facial Recognition: Cross-Modal Experiments
Combined language embeddings from Google Word2Vec with vision embeddings from the UTK facial dataset to explore multimodal similarity spaces and emergent associations between image and text domains.
Word2VecComputer visionPython
Toxicity Analysis on Wikipedia Talk Pages
Data analysis and inferential statistics on a 127,820-comment dataset produced by a natural language toxicity classifier, including error analysis on misclassified toxic and non-toxic comments.
PythonStatistics
Fairness Analysis of Deaths in Custody, California
Statistical and fairness analysis in Python of deaths that occur in custody or during arrest in California, using the state Department of Justice open data.
PythonPandas
Ad Information Retrieval with Transformers
A transformer-based retrieval pipeline that classifies personal ad-targeting data into categories such as technology, education, health and housing, as an exercise in applying pretrained language models to personal data analysis.
TransformersPython
GameSwap: Peer-to-Peer Game Bartering Platform
Built across three phases from EER and information-flow diagrams to a relational schema to a working application: MySQL backend, Angular and Node.js frontend, Spring Boot and Java application layer, with black-box and white-box testing throughout.
JavaSpring BootAngularMySQL
JobCompare: Weighted Job Offer Comparison App
An Android application for multi-criteria comparison of job offers. Owned the UML class diagrams, deployment and component architecture, and the wireframes for the full user journey, with SQLite persistence and agile delivery.
JavaAndroidSQLite
Penetration Testing Campaign with GoPhish
A controlled phishing campaign deployed on a hardened Ubuntu VPS with domain hosting, SSL provisioning and authenticated SMTP relay, covering campaign design, landing pages, credential-capture analysis and reporting against authorised targets.
GoPhishLinuxSSL/TLSSMTP