2026-10-05
| Day | Topics | Slides |
|---|---|---|
| 1 | Setup and help - packages and environments - NumPy - SciPy | Day 1 |
| 2 | pandas - Matplotlib - scikit-learn - where to go next | Day 2 |
Each day has a break halfway through. Every block starts with a few slides, then exercises.
conda with conda-forge as the default channel.git clone https://github.com/neuroinformatics-unit/course-intro-scientific-python.gitEvery line should say OK.
In the repository
notebooks/: the exercisessolutions/: worked solutionsdata/: every dataset we use, so nothing downloads during classIn class
Material from these sources is adapted under their licences. See each notebook’s header for details.
While we get started: run jupyter lab, open notebooks/day1_01_getting_help.ipynb and run the first cell, the setup check.
| Day | Before the break | After the break |
|---|---|---|
| 1 | Setup and help - packages and environments - NumPy | SciPy |
| 2 | pandas - Matplotlib | scikit-learn - where to go next |
notebooks/, numbered by day and block# Your code here or Your answer here are for you--------------------------------------------------------------------------- TypeError Traceback (most recent call last) Cell In[1], line 4 1 import numpy as np 2 3 n_points = 100 / 8 ----> 4 x = np.linspace(0, 1, n_points) File ~/work/course-intro-scientific-python/course-intro-scientific-python/.venv/lib/python3.13/site-packages/numpy/_core/function_base.py:123, in linspace(start, stop, num, endpoint, retstep, dtype, axis, device) 27 @array_function_dispatch(_linspace_dispatcher) 28 def linspace(start, stop, num=50, endpoint=True, retstep=False, dtype=None, 29 axis=0, *, device=None): 30 """ 31 Return evenly spaced numbers over a specified interval. 32 (...) 121 122 """ --> 123 num = operator.index(num) 124 if num < 0: 125 raise ValueError( 126 f"Number of samples, {num}, must be non-negative." 127 ) TypeError: 'float' object cannot be interpreted as an integer
--------------------------------------------------------------------------- ZeroDivisionError Traceback (most recent call last) Cell In[2], line 7 3 4 def call_func(x): 5 y = divide_0(x) 6 ----> 7 z = call_func(10) Cell In[2], line 5, in call_func(x) 4 def call_func(x): ----> 5 y = divide_0(x) Cell In[2], line 2, in divide_0(x) 1 def divide_0(x): ----> 2 return x / 0 ZeroDivisionError: division by zero
Julia Evans, Debugging manifesto.
See also SPL, Getting help and finding documentation.
np.linspace with help(), ? and the online docsnp.lin<Tab>Open notebooks/day1_01_getting_help.ipynb
OKnp.arange three ways and answer the questionsStretch: section 4, tab completion and searching
pip installs packages from PyPI, the official Python package indexconda installs packages from conda or conda-forge, a community-maintained channeluv installs packages from PyPI and manages whole projects (uv pip is a drop-in for pip)numpy<2, project B needs numpy>=2.3
xkcd 1987, CC BY-NC 2.5.
./.venv)venv + pip, or uv~/miniforge3/envs/)condascientific-python is a conda environmentwhich python (macOS/Linux) or where python (Windows)| conda | pip | uv | |
|---|---|---|---|
| Install | conda install numpy |
pip install numpy |
uv add numpy |
| Pin a version | conda install tqdm=4.70.1 |
pip install tqdm==4.70.1 |
uv add tqdm==4.70.1 |
| Update | conda update scipy |
pip install -U scipy |
uv lock --upgrade-package scipy, then uv sync |
| List | conda list |
pip list |
uv pip list |
| Remove | conda remove numpy |
pip uninstall numpy |
uv remove numpy |
= for conda, == for pip and uvconda env list shows every environment on your machine; the conda cheat sheet has the restTerminal tips
Windows: use the Miniforge Prompt. --help on any command; Tab completes names; Up/Down recall history; Ctrl+C cancels; q quits a long output.
numpy=2.5uv writes one automatically: uv.lockconda needs an extra tool (conda-lock)Open notebooks/day1_02_environments.ipynb. The work happens in a terminal.
.venv, rebuild it from the lockfileWarning
Leave the course environment scientific-python alone: only work on the throwaway ones.
After SPL, Python scientific computing ecosystem.
array([0, 1, 2, 3, 4, 5])
array([[ 0, 1, 2, 3],
[ 4, 5, 6, 7],
[ 8, 9, 10, 11]])
reshape gives the same values a new shapeint64, float64, uint8, bool, …dtype=.astype()start:stop:step, stop excluded, counting from 0a[row, column], and : means “all”array([[99, 1, 2, 3],
[ 4, 5, 6, 7],
[ 8, 9, 10, 11]])
b looks at the same memory as anp.shares_memory(a, b)a[...] and += change it in place| Array | Shape | Axes |
|---|---|---|
lfp |
(5000, 384) | time (500 Hz, 10 s), channel |
spike_counts |
(51, 533, 30) | unit, trial, 50 ms bin |
contrast, choice, rt, … |
(533,) | trial |
area, depth_um |
(384,) | channel |
International Brain Laboratory, Brain Wide Map, via DANDI (International Brain Laboratory et al. 2026).
Open notebooks/day1_03_numpy.ipynb, Part 1: indexing
Stretch (section 5): views and copies, a function that changes its input by accident, int16 pitfalls
+, -, *, /, **np.sqrt, np.exp, np.log, np.sin, np.abs, …<, ==, … give bool arraysaxismean, std, min, max, argmin, argmax, any, all, cumsumnp.nanmean and friends ignore missing values (NaN)argmax: which student was bestkeepdims=True to keep the shapeArrays of different shapes combine by stretching axes of size 1.
Rules, comparing shapes from the right:
(3, 1) with (4,) → (3, 4) ✓(3, 2) with (3,) → ✗ errorarray([[ 0, 1, 2, 3],
[10, 11, 12, 13],
[20, 21, 22, 23]])
np.newaxis (or None) adds an axis of size 1Open notebooks/day1_03_numpy.ipynb, Part 2: aggregations and broadcasting
Stretch: distances between recording sites
bool array: a maskarray([False, True, False, False, True])
&, |, ~ (not and/or), with brackets around each comparisonnp.whereany and all%timeit: about 70 ms for a million values%timeit for one line, %%timeit for a whole cell in Jupyter or IPythonPDSH, Profiling and Timing Code. Timings measured on an Apple-silicon laptop; yours will differ.
Open notebooks/day1_03_numpy.ipynb, Part 3: masks and vectorising
NaNStretch: fancy indexing, data statistics (after SPL Ex. 23)

fig, ax = plt.subplots(): a figure with one set of axesax.plot(x, y) draws a line; "o-" adds markersOpen notebooks/day1_03_numpy.ipynb, Part 4: check your work with a plot
| Submodule | For |
|---|---|
scipy.signal |
Filtering, spectra, peak finding |
scipy.fft |
Fourier transforms |
scipy.stats |
Distributions, statistical tests |
scipy.optimize |
Fitting (curve_fit), minimising, root finding |
scipy.ndimage |
N-D image filters, morphology, measurements |
import scipy as sp, then sp.signal.butter, sp.stats.ttest_ind, …interpolate, integrate, linalg, sparse, spatialsp.fft has the forward and inverse transforms, in 1-D, 2-D and N-DOpen notebooks/day1_04_scipy.ipynb, Part 1: filtering with the FFT (after SPL Ex. 42)
scipy.signalimport scipy as sp
fs = 500
t = np.arange(0, 2, 1 / fs)
rng = np.random.default_rng(0)
noisy = np.sin(2 * np.pi * 3 * t) + rng.normal(0, 0.5, t.size)
sos = sp.signal.butter(4, 10, btype="lowpass",
fs=fs, output="sos")
smooth = sp.signal.sosfiltfilt(sos, noisy)
fig, ax = plt.subplots(figsize=(5, 2.5))
ax.plot(t, noisy, color="lightgrey")
ax.plot(t, smooth);
butter, iirnotch, … (cut-offs in Hz with fs=)sosfiltfilt (or filtfilt for b, a)filtfilt runs forwards then backwards: no delayaxis= filters every channel at onceOpen notebooks/day1_04_scipy.ipynb, Part 2: filtering with scipy.signal
cleanedStretch: a smoother spectrum with welch, which distribution fits the response times (Ex. 41), 2-D minimisation (Ex. 40)
Exercises adapted from Scientific Python Lectures (Varoquaux et al. 2024).
Exercises adapted from Scientific Python Lectures (Varoquaux et al. 2024) (CC BY 4.0). Further reading in the Python Data Science Handbook (VanderPlas 2016).
| gdpPercap_1952 | gdpPercap_1957 | gdpPercap_1962 | gdpPercap_1967 | gdpPercap_1972 | gdpPercap_1977 | gdpPercap_1982 | gdpPercap_1987 | gdpPercap_1992 | gdpPercap_1997 | gdpPercap_2002 | gdpPercap_2007 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| country | ||||||||||||
| Albania | 1601.056136 | 1942.284244 | 2312.888958 | 2760.196931 | 3313.422188 | 3533.003910 | 3630.880722 | 3738.932735 | 2497.437901 | 3193.054604 | 4604.211737 | 5937.029526 |
| Austria | 6137.076492 | 8842.598030 | 10750.721110 | 12834.602400 | 16661.625600 | 19749.422300 | 21597.083620 | 23687.826070 | 27042.018680 | 29095.920660 | 32417.607690 | 36126.492700 |
| Belgium | 8343.105127 | 9714.960623 | 10991.206760 | 13149.041190 | 16672.143560 | 19117.974480 | 20979.845890 | 22525.563080 | 25575.570690 | 27561.196630 | 30485.883750 | 33692.605080 |
| Bosnia and Herzegovina | 973.533195 | 1353.989176 | 1709.683679 | 2172.352423 | 2860.169750 | 3528.481305 | 4126.613157 | 4314.114757 | 2546.781445 | 4766.355904 | 6018.975239 | 7446.298803 |
| Bulgaria | 2444.286648 | 3008.670727 | 4254.337839 | 5577.002800 | 6597.494398 | 7612.240438 | 8224.191647 | 8239.854824 | 6302.623438 | 5970.388760 | 7696.777725 | 10680.792820 |
index_col: the column that becomes the row labelspd.read_excel, data.to_csv(...) for other formats and for writingdata.head(): the first few rowsdata.info(): columns, dtypes, missing valuesdata.describe(): summary statistics per columndata.shape, data.columns, data.indexdata.T: rows and columns swappedhead, info, describeiloc and loc| gdpPercap_1952 | gdpPercap_1957 | |
|---|---|---|
| country | ||
| Albania | 1601.0 | 1942.0 |
| Austria | 6137.0 | 8843.0 |
data["gdpPercap_2007"] is the shortcut for one column: means “all rows” or “all columns”Open notebooks/day2_01_pandas.ipynb. The exercises follow the Software Carpentry gapminder lesson (Software Carpentry 2024), on the recording’s trials and units tables.
idxmin/idxmax, selection practice| gdpPercap_1952 | gdpPercap_2007 | |
|---|---|---|
| country | ||
| Ireland | 5210.0 | 40676.0 |
| Norway | 10095.0 | 49357.0 |
&, | and ~, with brackets around each comparisonrich.sum() counts the matching rows; data[rich] keeps themmean, count, idxmax, …continent
Africa 54.8
Americas 73.6
Asia 70.7
Europe 77.6
Oceania 80.7
Name: lifeExp_2007, dtype: float64
Beyond today: merging tables is stretch material in the notebook; PDSH chapter 3 also covers missing data.
.aggOpen notebooks/day2_01_pandas.ipynb
Stretch: one grouping question of your own
| continent | country | gdpPercap_1952 | gdpPercap_1957 | gdpPercap_1962 | |
|---|---|---|---|---|---|
| 0 | Africa | Algeria | 2449.0 | 3014.0 | 2551.0 |
| 1 | Africa | Angola | 3521.0 | 3828.0 | 4269.0 |
| 2 | Africa | Benin | 1063.0 | 960.0 | 949.0 |
Open notebooks/day2_01_pandas.ipynb, Part 4: tidy data
Stretch: Many Ways of Access; merging units with electrodes (PDSH)

pyplot versus fig, axfig, ax yesterday; we use it throughout: it scales to many panels and works inside functionsplt.xlabel(...) calls become ax.set_xlabel(...)Matplotlib docs, the two interfaces.
bbox_inches="tight" trims the white borderpyplot style, then translate it back to fig, axOpen notebooks/day2_02_matplotlib.ipynb, Part 1: simple plot
Optional: limits, spines and an annotation
| Method | For |
|---|---|
ax.scatter(x, y, s=..., c=...) |
Points, with size and colour per point |
ax.hist(values, bins=...) |
Distributions |
ax.bar(names, heights) |
Values per category |
ax.errorbar(x, y, yerr=...) |
Means with error bars |
ax.imshow(array) |
Images and 2-D arrays |
fig, axes = plt.subplots(1, 3), then axes[0].plot(...)
imshow maps values through a colormapviridis, gray); diverging maps (RdBu_r) only for data centred on zerofig.colorbar(im, ax=ax) shows the mappingOpen notebooks/day2_02_matplotlib.ipynb, Part 2: the whole recording as an image (after SPL Ex. 31)
Stretch: a scatter of the units along the probe, spike counts as an image, the population response, subplot_mosaic layouts, the LFP’s envelope, more plot types
y from X (classification, regression)X alone (clustering, dimensionality reduction)X and yX: a table, one row per sample, one column per feature
trials[["contrast", "probability_left"]]y: a 1-D array, one answer per sample
X must be 2-D even with one feature: trials[["contrast"]], not trials["contrast"]array([False, True])
fit(X, y) learns from the datapredict(X) for new samplespredict_proba(X): one column per class, in the order of model.classes__: model.coef_cross_val_scorecoef_, then predict_proba on a grid, plotted over the dataOpen notebooks/day2_04_sklearn.ipynb
Stretch: a better feature, whether the block helps (cross-validation), decoding the choice from the neurons
import skimage as ski
image = ski.io.imread("nuclei.tif") # a NumPy array
smooth = ski.filters.gaussian(image, sigma=2)
mask = smooth > ski.filters.threshold_otsu(smooth) # a boolean mask
labels = ski.measure.label(mask) # 0 = background, 1..N = objects
table = ski.measure.regionprops_table(labels, properties=["area", "eccentricity"])import xarray as xr
# lfp_uv: (n_samples, n_channels); time: seconds; area: one label per channel
lfp = xr.DataArray(lfp_uv, dims=("time", "channel"),
coords={"time": time, "area": ("channel", area)})
lfp.sel(time=slice(2, 4)).mean(dim="time") # by label and by name, not position
lfp.groupby("area").mean()groupby on N-D data| Package | What it is for |
|---|---|
| Zarr | Chunked, compressed N-D arrays on disk or in the cloud |
| Polars | Fast DataFrames with a lazy query engine |
| Numba | Compiles the Python loops you cannot avoid to machine code with @numba.njit |
loc versus iloc, masks, groupby, and tidy (long) datafig, ax = plt.subplots()X is samples × features; fit, then score on held-out data against a baselineKeep going: the Scientific Python Development Guide, the advanced chapters of Scientific Python Lectures (Varoquaux et al. 2024), and the BioImage Analysis Notebooks.
Exercises adapted from Scientific Python Lectures (Varoquaux et al. 2024) (CC BY 4.0) and the Software Carpentry gapminder lesson (Software Carpentry 2024) (CC BY 4.0). Further reading in the Python Data Science Handbook (VanderPlas 2016).