Skip to content

Research Capabilities

My work combines experimental biological reasoning with computational analysis and scientific software. I move between molecular mechanisms, genomic data, and model behavior, carrying the biological question through to an interpretable result.

Biological questions into computational studies

I bring experience in molecular genetics, RNA biology, genome engineering, and computational genomics to the design and interpretation of analyses. My work includes microbial genome assembly, transcriptomic reanalysis, and research on how annotation choices affect measurements.

Research examples: genome assembly and comparative analysis · experimental genetics.

Biological sequence models and evaluation

MutScan evaluates existing ESM-family protein language models for zero-shot variant scoring. I develop comparisons that examine the information in model scores, their sensitivity to model checkpoints, and their relationship to gene-specific evidence.

  • ESM1b, ESM1v, ESM2 through 15B parameters, ESMC, and ESM3.
  • Site-level and substitution-type baselines, residual analysis, and classical substitution matrices.
  • Per-protein ROC/AUROC evaluation, stratified benchmarks, and accuracy-versus-compute comparisons.
  • Benchmark circularity, evidence dependence, and explicit failure analysis.
  • ClinVar, dbSNP, Ensembl, HGVS, MANE Select, AlphaFold pLDDT, and structural context.

Research example: MutScan model-audit methods and limitations. This work evaluates existing foundation models; it does not claim foundation-model training.

LLM-assisted scientific workflows

I use Claude and other LLM-based assistants to develop scientific software, explore analytical alternatives, and prepare figures and manuscripts. In MutScan, this spans a protein-general scoring pipeline and its downstream comparative analyses.

My role is to define the biological problem, review analytical assumptions, and check the resulting code and claims against data. Configuration checks, run provenance, baseline comparisons, and figure-to-table audits make that oversight concrete.

Research example: LLM-assisted development and analysis in MutScan. The project repositories remain private while the manuscript is in preparation.

Computational genomics and bioinformatics

  • Genome and transcriptome assembly, annotation, and comparative genomics.
  • Genomic, epigenomic, transcriptomic, proteomic, and metagenomic data.
  • Illumina, PacBio, and Ion Torrent sequencing.
  • Bulk and single-cell assays, differential expression, pathway enrichment, motif discovery, and regulatory-network analysis.

Research examples: published genomics studies and transcriptomic reanalysis. Student projects extend the lab's work into isoforms, public-data quality, and comparative proteomics.

Reproducible workflows and computing

  • Python, R, Bash/Shell, pandas, NumPy, scikit-learn, PyTorch, and Bioconductor.
  • Linux, HPC clusters, Slurm, parallel execution, and cloud computing.
  • CPU/GPU workflows, Git/GitHub, regression tests, and run-level provenance.
  • Documented tools and publication-quality reporting.

Public examples: released software · manuscript-template code and build documentation.

Experimental biology and collaboration

My experimental foundation includes gene regulation, epigenetics, RNA silencing, mutant generation, and DNA, RNA, and protein methods in bacterial and fungal systems. That background helps me communicate computational findings in terms of biological mechanisms, controls, and testable hypotheses.

I lead interdisciplinary research, mentor graduate and undergraduate scientists, and collaborate across experimental and computational biology. In Digital Biology, students practice documented computation and verify AI-assisted work against biological reasoning and reproducible evidence.

Selected research Discuss collaboration