Founding Engineer

SoulBio

I joined SoulBio early and worked across data platforms, scientific applications, and client projects. As the company grew, I took on more product and technical direction. That includes deciding which problems to work on, where AI belongs in a scientific workflow, and which ideas are worth building.

AI and technical direction

  • Defined SoulBio’s Scientific AI Enablement offering from the ground up: the problems it should solve, the core capabilities it needed, how knowledge and workflows should fit together, and how the resulting systems should be evaluated.
  • Designed a biotech context graph that connects scientific knowledge with internal decisions, documents, and organizational context, then built a working proof of concept for customer discussions.

Scientific platforms and client work

  • Built and deployed drug targetability and selectivity models that reduced hypothesis-validation time by ~40% for oncology teams.
  • Architected and built ingestion and computation workflows for bulk and single-cell RNA-seq on a ~1 TB PostgreSQL system, bringing runs down from close to a day to a few hours.
  • Built RNA-seq pipelines, APIs, analysis and visualization tools, and CI/CD workflows used in day-to-day scientific work.
  • Managed AWS infrastructure for deployed applications, including reliability, scaling, performance, and cost trade-offs.
  • Wrote product requirements and technical specifications for internal platforms and new client projects.

Applied ML and research

  • Co-developed an LLM-powered search system that lets researchers query a corpus of 250K+ GEO datasets in natural language.
  • Built a bioinformatics analysis agent for RNA-seq that coordinates multi-step workflows, automates repetitive analysis, and leaves scientific interpretation and decisions with the researcher.
  • Co-authored research on the limits of GPT-based cell-type annotation and contributed technical articles and whitepapers.

Data Analyst

Elucidata

  • Curated large-scale biomedical datasets using data-centric models that supported FAIR data generation and machine-learning research.
  • Implemented information-extraction models and pipelines for GEO and PubMed, improving accuracy by about 20%.
  • Co-authored a multi-task learning approach that made biomedical NER inference three times faster.
  • Fine-tuned a proprietary 3B-parameter language model using PEFT, improving performance by 40% on biomedical curation tasks.
  • Built a GPT-assisted ontology-normalization tool with 96% mapping accuracy.

Python Developer Intern

MindBowser

  • Built and tested REST APIs in Python.
  • Improved data-processing scripts through refactoring and modularization.