AI agents in Coding and Science


Richard Neher
Biozentrum & SIB, University of Basel


slides at neherlab.org/202609_facultylunch.html

Chatbot

Chatbot: closed loop between user and model
e.g. ChatGPT, Claude.ai, Gemini

Agent

Agent: model uses user resources, checks and iterates
e.g. Claude Code, Codex, Cursor
Same model – but different tools and feedback cycles

Why would you want to do this?

  • Longer independent workflows: hand over a task, not a question
  • Feedback comes from the environment, not from you
    • Coding: build, run the tests, read the errors, fix – the test suite is the harness
    • Science: run the analysis, inspect the output (plots, summary statistics, sanity checks), revise
  • You review results and key decisions, not every step

The longer the leash the higher the risks

  • Unintended actions
    • deleted or overwritten files and data
    • leaked credentials or unpublished data
    • rogue actions on shared resources: repos, clusters, email
  • Wrong results
    • plausible output that is wrong
    • insufficient tests, or the agent adapts the tests until they pass
    • more output than you can verify
  • Code as liability
    • large, complex code base that nobody understands well enough to maintain
    • bugs surface later, with no one who knows where to look

Mitigation: contain, verify, simplify

  • Contain
    • sandbox: agent writes only to the project directory
    • no credentials or sensitive data within reach, restricted network
    • everything recoverable: version control and backups
  • Verify
    • checks you define: known results, positive controls, tests you have read
    • inspect intermediate outputs, not just the summary
    • scope tasks to what you can verify
  • Simplify
    • ask for less code, reuse established libraries
    • review diffs, insist on simple architecture
    • treat it as code you are responsible for

Example: fitness effects of Spike mutations on structure

Build a static website (TypeScript/React) that shows the fitness effects of amino-acid substitutions in the SARS-CoV-2 Spike protein, using data from our project with Jesse Bloom: jbloomlab.github.io/SARS2-mut-fitness/S.html

  • Map the average fitness cost of mutations at each site onto a structure of the Spike trimer.
  • Effects are estimated relative to different wildtype sequences (clade founders); let the user choose among them.

Before writing any code, propose a plan: data sources and preprocessing, structure and viewer, and how you will check the result. Wait for my OK.