Using Natural Language to Do Science

How Agentic AI Empowers Astronomers

James Nightingale

By the end of this workshop

  • Natural language is now a way to do science, not just to ask questions.
  • Agentic AI: how Claude Code and Codex differ from ChatGPT.
  • The incredibly powerful things it can do.
  • Where it can go wrong.
  • What happens next: 25 Claude subscriptions for scientists.

For the majority of this talk I will drink the AI kool-aid.

At the end of the workshop we will discuss the pitfalls and challenges surrounding AI.

Remember This?

  • You spot a tiny issue with a matplotlib figure.

Remember This?

  • You spot a tiny issue with a matplotlib figure.
  • You change the python code, but it doesn't change the figure.

Remember This?

  • You spot a tiny issue with a matplotlib figure.
  • You change the python code, but it doesn't change the figure.
  • You trial-and-error StackOverflow answers for the next four hours.

Remember This?

You wish you could just tell matplotlib what you want the figure to look like.

Remember This?

Remember This?

We don't have this problem anymore.

AI lets us use natural language to interface with matplotlib.

Nobody became an astronomer to learn matplotlib syntax.

[Is this controversial? Who disagrees?]

What Is Agentic AI?

A chatbot answers. An agent acts.

[Live demo: the same request to ChatGPT, then to Claude Code.]

The tools

Claude Code (Anthropic) and Codex (OpenAI) run in a terminal, inside your project folder.

They can read every file, run any command, edit code and check their own work, and they ask before doing anything that cannot be undone.

You talk to them the way you would brief a capable student.

Gemini CLI, Cursor and Copilot are the same idea with different wrapping.

Example 1: One prompt to Claude

The folders @PyAutoGalaxy and @PyAutoLens contain the PyAutoLens source code. Find the power-law mass model used by PyAutoLens.

Then, compare its parameterization to the COOLEST lens model standard: https://github.com/aymgal/COOLEST

Then, on the PyAutoLens GitHub repository, put up an issue comparing the two parameterizations and how they differ.

What came back

A GitHub issue, written and opened by the agent.

It read both code bases and the COOLEST package, and wrote out both convergence equations.

github.com/PyAutoLabs/PyAutoLens/issues/563

The part that matters

Copying the Einstein radius straight into COOLEST gives a q-biased model. The agent found that, and said so.

Example 2: Paper LaTeX

No More LaTeX!

I have lost many evenings trying to make a LaTeX table in a paper look right.

One prompt to Codex

The output folder contains lens modeling results for 3 SLACS lenses.

Make a paper folder with a main.tex using the MNRAS template. Then, make a latex table containing the name and Einstein radii of each lens, including 3 sigma errors, using the model.results file in each lens's results folder.

Add a table caption describing its contents.

No More LaTeX!

Now I tell Codex which results go in the table and what it should look like. It pretty much does it.

A paper is a folder

One .tex file per section, the plots, the outputs, and a .agents folder that tells the agent how the paper is written.

The agent handles paper formatting, referencing and building the PDF. I read, argue and decide.

[Live demo of a paper folder: how the whole writing process uses AI now.]

Nobody became an astronomer to learn LaTeX syntax.

[Still not controversial?]

Example 3: HPC

Two weeks before the papers were due

A collaborator posts a new catalogue. Every lens has been re-ranked. Every model I had run now has to be redone.

What this used to look like

ssh in, sbatch, squeue, tail -f the log, rsync the results down. Repeat until the postdoc is over.

So I pasted the Slack message into Claude

I then guided it through reorganising the results on the HPC.

Nobody became an astronomer to learn ssh, rsync and Slurm syntax.

[Surely we can all agree on this one?]

The syntax was never the science

Turning an idea into an analysis has always meant translating it through matplotlib, LaTeX, Slurm and Python. That syntax was only ever the interface between your thinking and the computer.

Now the interface is natural language. You say what you want in the words you already think in, the computer does it, and you keep the understanding, direction and judgement.

Building Knowledge In a Wiki

Knowledge that grows

Everything the agent learns, from papers, results and conversations, is written into a wiki in natural language.

Each new task builds on that knowledge, so the agent knows more about your science every time you use it.

Using Natural Language to Do Science

PyAutoLens AI Assistant

The COSMOS-Web Ring, imaged by JWST.

Your starting promptCopy
I want to use the PyAutoLens Assistant: https://github.com/PyAutoLabs/autolens_assistant First clone that repository, cd into it and follow its AGENTS.md.

I'd like to understand how gravitational lensing works using the JWST image of the COSMOS-Web Ring that ships with the assistant. Show me the picture, explain what we are looking at, and walk me through fitting a lens model so we can measure the mass inside the ring and see how well the model reproduces the observations. Pitch it at my level: ask me what my background is first. Explain what we are doing as we go, and let me ask questions or change the analysis along the way.

[Live demo: paste this into Claude Code or Codex, then go to the autolens_assistant GitHub repository.]

3 More Really Powerful Use Cases

A new sampler in an hour

  • Handed the paper, the agent implemented a whole new sampling method in PyAutoLens with minimal human oversight.
  • The tests are what made it safe to accept.
  • I now have an inference wiki that continually runs new samplers on PyAutoLens likelihood functions.

Vibe coding can be great

The COSMOS-Web Lens Survey gallery: every candidate lens's JWST image, metadata, PyAutoLens results and .fits downloads, one click away.

I described the site I wanted and the agent built it. I never read the HTML or JavaScript.

COWLS lens gallery and GitHub repository

[Live demo: the COWLS gallery and GitHub repository.]

Natural-language software development

Humans describe what the software should do, why it is needed, and how success will be judged.

Specialist agents plan, implement, test and prepare releases.

Humans remain responsible for scientific objectives, contributor communication and every consequential decision.

[Live demo: the PyAutoScientist dashboard. github.com/PyAutoLabs/PyAutoScientist]

AI Pitfalls

When I started using AI, I had used matplotlib, LaTeX and HPCs for over ten years. I wrote PyAutoLens. I have done lens modeling my whole career.

I learned the statistics and the software development myself, the hard way.

AI is incredibly powerful when you already understand the core concepts of what you are doing.

Learning the science

Weighing a galaxy in natural language does not remove the need to understand the physics, the model assumptions and the uncertainties.

Easier access to sophisticated analysis makes understanding more important, not less.

So I build the teaching in: HowToLens teaches lensing from first principles. Learn first, then put that knowledge into practice.

[Live demo: a HowToLens lecture in Google Colab.]

Other challenges

  • Data privacy. Your data, code and ideas go to a giant tech company. Do we want them owning the world?
  • Environmental impact. GPUs use a lot of energy, and every prompt spends some of it.
  • Job automation. If the agent does the work, what is the role of a PhD student, a postdoc or an astronomer?

Claude Team for Scientists

We have 25 free Claude subscriptions for the year.

I am keen to lead the rollout, but we should agree on how we approach it first.

A chatbot answers. An agent acts: it has tools, and it loops. The loop is the whole difference. It sees the error you would have pasted, and the file you would have opened, without you in the middle.

Three sentences. Two code bases, an external standard and GitHub. Nothing below this was typed by a human.

Follows Karpathy's "LLM wiki" idea: a curated wiki the model reads and extends, rather than retrieval over raw PDFs.