pymodelvis · neural_flow for PyTorch

See what your network sees.

Graph viewers draw a model's operations. neural_flow draws what one specific input becomes inside it — stage by stage, from pixels or voxels to the prediction — with the evidence for each step. As figures for papers and talks, an interactive explorer, and movies over changing inputs.

$ pip install "pymodelvis[all] @ git+https://github.com/MASILab/pymodelvis"
ResNet-50 activation flow for a photograph of a cat, from the input through feature-map stacks to the prediction 'Egyptian cat'
ResNet-50 looks at a cat. Stacks are the most active feature maps at each stage (front page: all channels as colour). Beams show which region of one stage feeds the next; circles show what a unit sees in the photograph; lines into the head are weight × activation; the heat map is Grad-CAM.

One command, no code

Installing adds a neural-flow command. Point it at a model and an input; it picks the stages worth showing, runs the model once, and writes the figure. Models can be built-in names, any torchvision or timm model, a MONAI bundle, a saved model.pt, or your own my_net.py:Class with a checkpoint.

Command-line guide →

terminal
# the cat above
$ neural-flow demo cat

# your photo, a real model
$ neural-flow render resnet50 -i photo.jpg \
      --style cinematic -o flow.png

# your 3-D U-Net on a NIfTI volume
$ neural-flow render my_unet.py:UNet3D \
      --weights best.pt -i t1.nii.gz --crop 96

# a movie: which region matters?
$ neural-flow movie cxr --occlusion chest.png -o occ.mp4

…or one line of Python

visualize_model works with any nn.Module: CNNs, U-Nets, vision transformers, 3-D medical networks, and models with several inputs or heads. Hooks are always removed and the model is never modified.

Python guide →   API →

python
from neural_flow import visualize_model, animate_inputs

visualize_model(model, x, output="flow.png",
                style="cinematic")
visualize_model(model, x, output="explore.html")

# several inputs, several heads, 3-D volumes
visualize_model(model, {"mri": vol, "clinical": v})

# a movie over a sequence of inputs
animate_inputs(model, frames, output="pan.mp4")

Designed to be trusted

Automatic

Picks the 5–12 stages worth showing from arbitrary architectures and folds away activations, normalization and reshapes. Override any choice by name, type or regex.

Faithful

Every pixel comes from the actual input. Outputs are shown as probabilities only when they are probabilities, and summarized stages are marked ≈.

Explanatory

Receptive fields, stage-to-stage dependencies, weight × activation contributions and Grad-CAM, all computed from gradients for the input you give it.

3-D native

[B, C, X, Y, Z] is a first-class object: anatomy becomes translucent voxel blocks, with activation-driven orthogonal slices and 3-D segmentations.

Safe at scale

Large activations are summarized on the model's own device within a memory budget. Visualization problems never break your forward pass.

Presentation-grade

Black cinematic, technical and teaching styles; PNG, SVG and PDF; an interactive HTML explorer; MP4 and GIF movies; a bundled typeface.

From photographs to MRI volumes

Real pretrained weights where they exist; demo models trained on synthetic data otherwise.

Full gallery

Built for 3-D medical imaging

[B, C, X, Y, Z] volumes are first-class. Transformer U-Nets such as UNETR, Swin UNETR and MASI's UNesT show their transformer levels as stages and their decoder as a U; the output card shows the whole scan fused from sliding windows, with the traced window outlined; everything is drawn to scale from the voxel spacing.

UNETR activation flow on a 3-D head volume
UNETR · whole-head segmentationThe transformer runs on a 4 × 4 × 4 token grid; its taps bridge to the decoder levels. 3-D models →
nnU-Net activation flow on a CT, TotalSegmentator organ model
nnU-Net · TotalSegmentator on a CTA trained nnU-Net loaded from its results folder, with nnU-Net's own preprocessing and sliding windows. nnU-Net →
terminal
# MASI UNesT, 133 brain structures, whole brain from sliding windows
$ neural-flow fetch unest
$ neural-flow render unest -i T1_mni.nii.gz --sliding-window --style cinematic

# watch 3-D inference, window by window
$ neural-flow movie unest --inference T1_mni.nii.gz --max-windows 16

# any trained nnU-Net, or TotalSegmentator on its example CT
$ neural-flow render nnunet:$nnUNet_results/Dataset123_Liver -i case.nii.gz --sliding-window
$ neural-flow render totalseg -i sample:ct --sliding-window --style cinematic

Movies over changing inputs

Stages, channels, colours and scales are held fixed across frames, so everything that moves is the network responding.

Camera pan · ResNet-50A window slides across four photographs: cat → espresso → drilling platform → go-kart.
Ageing subject · multi-head 3-DA synthetic subject ages while a lesion grows. Brain age, lesion probability and the segmentation track it.
Occlusion · chest X-rayA grey patch slides over the radiograph. Cardiomegaly drops when the patch covers the heart.
Camera pan · ViT-B/16The same four photographs through a vision transformer, token grids and all.
3-D inference · UNETROne sliding window per frame: the stages follow the window through the head while the whole-head segmentation assembles.
3-D inference · nnU-Net on CTTotalSegmentator's organ model works through a CT window by window (27 windows of 128³) while the organs assemble.

How movies work →

One trace, several audiences

The same captured activations, drawn for a talk, a paper, a classroom or a poster.

style=
style="cinematic" — black, feature-map stacks, explanations

How it works

  1. Trace

    Forward hooks and a runtime dataflow tracer record every module call, its shape, and the real wiring: skips, merges and branches.

  2. Select

    A module-tree cut picks 5–12 stages that change the representation, grouped into concepts such as “low-level features” or “decoder”.

  3. Explain

    Gradients give each stage's receptive field, its dependency on the previous stage, head contributions and Grad-CAM.

  4. Render

    2-D maps, 3-D volumes, tokens, vectors and attention become pictures: figures, the HTML explorer, or movies.

Architecture →   What beams, circles and lines mean →

Interactive explorer interactive

Interactive explorer

One self-contained HTML file: click a stage to see its module, shape and value statistics, switch the channel ranking, expand individual channels and browse attention heads. No server needed.

Open the ResNet-50 explorer →

Slide deck preview .pptx

A slide deck, made from the outputs

A 15-slide PowerPoint with the figures, four embedded movies and speaker notes — an example of what the tool produces for a talk, and a template for your own.

Download the deck →

Open source, from the MASI Lab

Developed at the Medical-image Analysis and Statistical Interpretation (MASI) Lab and VALIANT, Vanderbilt University. Released under the BSD-3-Clause license. Issues and pull requests are welcome.

cite
@software{landman_pymodelvis_2026,
  author  = {Landman, Bennett},
  title   = {pymodelvis / neural_flow:
             representation-flow visualization
             for PyTorch models},
  year    = {2026},
  version = {0.1.0},
  url     = {https://github.com/MASILab/pymodelvis},
  license = {BSD-3-Clause}
}