Connecting data and tools across DISC and beyond

Ashley Smith 0000-0001-5198-9574
mag.earth Ltd
Swarm Data Quality Workshop
2026-10-08

Background

  • Products: data and models
    • Swarm+Multimission, ESA official (owned and distributed by ESA; created by DISC)
    • Swarm and friends, external
  • Services
    • VirES, HAPI, Timeline Viewer, …
  • Software
    • viresclient, SwarmPAL, ChaosMagPy, …
  • Docs & training
    • Handbook, Notebooks, …
  • Reassess what we have developed in the context of Swarm
  • Consolidate and extend for the future
  • What does “beyond DISC” mean?

Swarm Product Data Handbook catalogue page for SW_MAGx_LR_1B: description and data-access links (files, VirES, notebook, HAPI) VirES for Swarm web client: layer panel, 3D globe with Swarm Alpha magnetic-intensity tracks, latitude plots of F and model residual, parameter histograms and a time slider VirES for Swarm HAPI server landing page listing available datasets, SW_MAGx_LR_1B among them Swarm Notebooks page 'MAGxLR_1B (Magnetic field 1Hz)' showing viresclient code that fetches the data

Swarm Handbook · VirES web client · VirES HAPI · Swarm Notebooks

Challenges and opportunities as we grow

More missions, products, tools, teams, papers

  • Complexity grows combinatorially, not linearly
    • Multiplying connections to track and friction in collaboration
    • Discoverability gets harder - too many things to find!
    • A participant at EHC*: “Swarm data is difficult!”
      • With N missions and M products, it becomes NxM more difficult!
  • Increasing value in shared services and standardisation
    • Reinventing and maintaining new things gets expensive
    • Reduce the surface area for:
      • the user community (ease discoverability)
      • connecting with wider environment (see IHDEA, E-SWAN etc)
  • Connections and interactions create new possibilities we haven’t thought of yet

→ context is scattered across institutions: hard to answer questions that cross a boundary

Data and software: Current status

  • Growing range of software tools
    • How do we connect them and make them discoverable?
    • We need a software handbook as well as the data handbook?
  • Complexity of cataloguing, documentation, and data publishing (particularly with multimission)
    • How do we adapt experience from Swarm to do better with the newer missions?
    • We need:
      • Stronger standards on data format and use of verification tools
      • Automated checks (eg ongoing: validation of data on ingestion to VirES)
      • Data citation mechanism considered before publication
    • There is much value in legacy data - and difficulty in preserving and accessing it
      • Are we creating much bigger problems today for future scientists? What small things can we do to alleviate that?

Software in Swarm

  • Lessons from the Python in Heliophysics Community (PyHC):
    • Slow to evolve from “just a list of packages” to more:
      • Standards and procedure: Mostly adopted from broader communities (eg PyOpenSci)
      • Identified core packages, and some shared issues to tackle collectively (eg data models & adapters)
    • Real coordination is hard - but even just sharing approaches and learning from each other is valuable
    • People will continue to build targeted bespoke tools, but usually building as part of a bigger project (eg sunpy, plasmapy) will be longer-lived
  • Do we need a (Swarm) software community and what would it look like?
    • Bridging Heliophysics / Geospace and Solid Earth
    • Desilo efforts behind each mission / project
  • Online hackathon TBD

(some of the) Python ecosystem: package bubbles around a stack of example figures. DISC-affiliated packages (viresclient, swarmpal, chaosmagpy, geospacelab, pyamps, pyswipe, swarmface, ibpmodel) in red; other packages (sunpy, spacepy, pysat, pyspedas, plasmapy, kamodo, hapiclient, speasy, sciqlop) in blue. Only core PyHC packages shown.

Software development

  1. Decide what you want to build

  2. hammer Build it!

    1. Green field project: can go fast

    1. Discover solution space

      • Figure out what you really want
      • Figure out what is possible/feasible
      • Feedback: Can you just add this?
      • Fixing things
    2. “Finish” it (there is always more)

  3. rainbow People use it

→ Most time is spent on exploration, negotiation, and discovering requirements
→ Particularly true with research software

→ And usually you end up with something different from the original goal

How to draw a horse, by Van Oktop, steps 1 to 4: two circles, stick legs, a smiley face, a scribbled mane and tail

Step 5, 'Add small details': a fully shaded, realistic drawing of a horse How to draw an owl, step 2: a finely shaded pencil drawing of a great horned owl on a branch (the meme's 'draw the rest' panel, caption cropped off)

How to Draw a Horse, Van Oktop (meme) · How to Draw an Owl (meme)

On AI coding

  • The bad: quality, waste, verbose bullshit, externalities, ethics, cognitive effects…
    • You should be sceptical
    • AI usage is an amplifier, both good and bad
  • Contemporary models are very capable for small software tasks
    • The choice of harness is perhaps more important (eg Claude Code, OpenCode, …)
    • The more complex the task, the riskier it is to rely on

Modes of usage:

  • ✔ Assistant (code suggestions, scripts, pair programming)
  • ✔ Prototyper (full code generation, exploratory)
  • ? Documentation (good to automate parts; needs a lot of intervention to create good docs)
  • ? Autonomous developer
    • ✔ Automated bug reports and code review
    • ?✔ Small maintenance and bugfixes
    • ?✔ Hyper-personal software
    • ? Full project and feature building (requires skill: agentic engineering)
  • ✔ Disposable code ……… ? Durable code

Close-up of the Tyrell Corporation's artificial owl from Blade Runner, one eye lit with the red replicant eye-shine

Do you like our owl?

It’s artificial?

Of course it is.

Must be expensive.

Very. Blade Runner (1982)

AI is great for the first 90%:

AI speeds up at firststill yours

What if?… Geomagnetic model explorer

What if?… Cross-mission catalogue

What if?… Knowledge graph

Data / software network

“INTERMAGNET in space” / “SwarmNext” /…”

data and software environment (“SwarmNext” is placeholder)

  • Umbrella organisation outside of ESA?
    • Carry over experience and momentum of Swarm DISC; broaden scope
    • Coordinate “internally” between missions; unified interface with IAGA, IHDEA, COSPAR, …
    • Core of Swarm friends (LEO magnetometry) with outer shell of acquaintances (eg SMILE, Plasma Observatory, GDC, …)
  • Shared product harmonisation and validation systems
    • Tools/services to validate products
    • Adopt common standards and procedures to smooth interfaces
      • stronger standards on format and metadata: “Swarm-like format” is not enough!
  • Shared distribution of products where possible (i.e. VirES)
  • Shared tools for L2+ processing (enable R2O2R, i.e. SwarmPAL)

Summary and suggestions

  • Growing complexity motivates better, and inter-connected, data/software curation
    • Problems:
      • Data with different owners/publishers live in different places
      • Hard to navigate and discover what exists
      • Dependency chains break things
    • Solutions:
      • Standardised metadata to allow catalogues to communicate
      • Unifying interfaces over the top of those catalogues
      • Knowledge Graph / Graph database to track connections and dependencies?
      • Monitoring of processing chains - with notifications & issue tracking
  • More software is being built (and AI accelerates this - unsustainably?)
    • Need to catalogue and coordinate for greater return
  • Opportunities with AI
    • Rapid prototyping
    • Great for low-stakes code: eg educational tools, visualisation
    • More accessibility may result in an explosion of tools

Read more

AI-assisted coding

Knowledge graphs for heliophysics

Software catalogue

Metadata standards