Research lab
Golden Toad ResearchEst. 2026

Research on the reliability of intelligent systems.

We test whether the accepted ways of evaluating AI capture how these systems behave once they leave the benchmark, with current work in medical imaging and neuroscience.

Areas of inquiry05

Each project starts from one question: does the standard metric measure what we care about? Select a bubble to explore an area.

Featured study · Medical imaging01

A routine scanner setting can move AI tumor measurements past a clinical threshold.

Using the National Lung Screening Trial, we compared AI segmentation on the same CT acquisition reconstructed with two different filters. Overlap scores did not change. Measured tumor volume did.

Patients whose AI-measured volume changed by more than 25% between reconstructions
2.5D U-Net
24%
nnU-Net
15%
With one radiologist box
3%

Box-prompt result is exploratory. Protocol fixed before training; test set scored once.

Computational neuroscience02

I Have No Body, and I Must Steer

A spiking model of the fruit fly connectome, given a simple body and placed in a hide and seek game. Silencing two dozen neurons in its compass pathway leaves the seeker almost unable to win.

Publication pending
Neural signals03

EEG and cognitive performance

Brain wave classification and Hurst exponent analysis of long-range structure in EEG recordings.

Publication pending
Media forensics04

AI-generated video detection

A detector for AI-generated video that checks visual, temporal, and audio evidence separately, explains each flag, and attributes the clip to its likely generator.

Publication pending
The lab

A research lab for testing whether AI is measured the way it is used, across medical imaging, neuroscience, and media.

Read about the lab and how to collaborate.

About the lab
Follow the work

LinkedIn

Updates on publications and new projects.

linkedin.com/in/jamirwin
Collaboration

Open to research partnerships and inquiries.