Research

Making scholarship more reproducible, durable, and useful.

I study research infrastructure as a connected system—from computational workflows and data standards to repositories, publishing platforms, governance, and the communities responsible for them.

Research statement

The method does not end when the computation finishes.

Scientific knowledge depends on a chain of systems and decisions that is often treated as invisible: how data are described, how software is preserved, how workflows are recorded, how credit is assigned, and how research moves into the scholarly record.

My research makes that chain visible and then asks how it can be improved. I am interested in infrastructure that carries evidence and context across disciplinary and institutional boundaries—especially infrastructure that supports verification, responsible reuse, and long-term stewardship.

This work grows from my background in computational biology and software engineering. Problems I first encountered in genomics—fragmented data, unrecorded methods, fragile software environments, and disconnected publications—are now central to my work in research data and scholarly communication.

Dendrogram comparing thousands of metabarcoding datasets by environmental source
A comparative view of metabarcoding datasets—an early example of the scale, provenance, and reuse problems that continue to shape my infrastructure research.
Core themes

Four layers of the same problem.

Scholarly infrastructure

Repositories, publishing platforms, open-access models, preservation, and the institutional systems that sustain the scholarly record.

Reproducibility

Executable workflows, software environments, provenance, benchmarks, research objects, and links between computation and publication.

FAIR and AI-ready data

Biocuration, metadata, persistent identifiers, data relationships, quality, documentation, and responsible machine use.

Trust and security

Cyberbiosecurity, research integrity, controlled access, data provenance, and resilient systems for sensitive scientific work.

Community infrastructure

Standards work is research work.

Shared vocabularies, identifiers, and conventions are not administrative details. They determine whether data can be discovered, combined, interpreted, and trusted.

I contribute to community efforts in agricultural data, biodiversity genomics, scientific literature, and responsible data use. This includes work with AgBioData, i5k, the Research Data Alliance, BIO-ISAC, and community groups developing guidance for genome assembly nomenclature and FAIR scientific literature.

Selected publications and outputs

Recent work.

This is a selected list. The complete, maintained record is available through ORCID.

  1. 2026

    Toward standardization in arthropod and biodiversity genome projects

    GENETICS

    DOI
  2. 2025

    Guidelines for gene and genome assembly nomenclature

    GENETICS

    DOI
  3. 2025

    AgBioDatabase Finder: an online tool to help researchers find and submit agricultural genomic, genetic, and breeding data

    microPublication Biology

    DOI
  4. 2025

    System Security Assessment Plan Template

    Research Data Alliance Artificial Intelligence and Data Visitation Working Group

    Report
  5. 2024

    FAIR Header Reference genome: a TRUSTworthy standard

    Briefings in Bioinformatics

    DOI
  6. 2024

    otb: an automated HiC/HiFi pipeline assembles the Prosapia bicincta genome

    G3: Genes|Genomes|Genetics

    DOI
  7. 2022

    met v1: expanding on old estimations of biodiversity from eDNA with a new database framework

    Database

    DOI
From questions to systems

Interested in building something together?

I am available for collaborative projects, student supervision, advisory work, research-data and repository planning, standards development, and research infrastructure design.