Areas of inquiry
Golden Toad ResearchFive areas

Five areas, one question.

Across medicine, neuroscience, and media, we ask whether the way an AI system is measured matches the way it is used. Each area below outlines what we have worked on and where we hope to go next.

Select an area

Select a bubble to jump to that area.

01

Measurement validity

01
Overview

Whether the number used to approve an AI system tracks the decision people make with it.

Where we have worked
  • Compared standard overlap scores with the clinical quantity a radiologist acts on, under changes the model should not care about.
  • Built evaluation protocols that are fixed before training, with held-out data scored only once.
  • Studied cases where a system looks correct in its final answer while its underlying behavior is not.
Where we are heading
  • Stability checks for AI measurements across routine changes in how data is collected.
  • Practical reporting guidance that pairs headline scores with the quantity that drives decisions.
  • Reusable test suites that other groups can run on their own models.
02

Medical imaging

02
Overview

AI for lung cancer screening, and how its measurements hold up across real scanner settings.

Where we have worked
  • Segmented lung tumors in National Lung Screening Trial CT scans with several model families.
  • Measured how reconstruction filters chosen by each site change AI tumor volume.
  • Tested whether a single radiologist box prompt makes volume measurements more stable.
Where we are heading
  • Other acquisition settings such as dose and slice thickness.
  • Growth assessment across repeat scans of the same patient.
  • Multi-site validation with clinical partners.
03

Computational neuroscience

03
Overview

What a wiring diagram of a brain can and cannot explain once it is given a body.

Where we have worked
  • Ran a spiking model built from the fruit fly connectome as the controller of an agent in a game.
  • Showed that the fly steers with its own compass circuit, and that silencing that circuit breaks navigation.
  • Found behaviors the wiring diagram alone does not reproduce, which point to what is missing.
Where we are heading
  • Richer bodies and sensory input for connectome-driven agents.
  • Direct comparison with recorded behavior from real flies.
  • Other circuits, including ones involved in learning and memory.
04

Neural signal analysis

04
Overview

Reading cognitive state from brain activity with methods that respect the structure of the signal.

Where we have worked
  • Classified EEG recordings collected during cognitive tasks.
  • Used the Hurst exponent to describe long-range structure in brain wave signals.
  • Related signal features to measures of cognitive performance.
Where we are heading
  • Models that generalize across people, not only within one subject.
  • Real-time analysis suited to wearable recording.
  • Links between signal structure and attention or fatigue.
05

Media forensics

05
Overview

Detecting AI-generated video and explaining why a clip was flagged.

Where we have worked
  • Built a multi-agent system that checks visual, temporal, and audio evidence separately.
  • Produced explanations for each flag so a reviewer can see where the evidence came from.
  • Added attribution that estimates which generator likely produced a clip.
Where we are heading
  • Keeping pace with new video generators as they are released.
  • Calibrated confidence that holds up on compressed, real-world footage.
  • Provenance tools that work alongside detection.
Collaboration

Working in one of these areas? We would like to hear from you.