Data loading and saving

Use these guides to load experimental data, add missing acquisition metadata, find scans, and save results.

Loading data from a supported endstation

Select the loader for your endstation:

Loader

Description

Loader class

da30

Scienta Omicron DA30 with SES

erlab.io.plugins.da30.DA30Loader

erpes

KAIST home lab setup

erlab.io.plugins.erpes.ERPESLoader

esm

NSLS-II Beamline ID21 ESM

erlab.io.plugins.esm.ESMLoader

hers

ALS Beamline 10.0.1 HERS

erlab.io.plugins.hers.HERSLoader

i05

Diamond Beamline I05

erlab.io.plugins.i05.I05Loader

kriss

KRISS ARPES-MBE

erlab.io.plugins.kriss.KRISSLoader

lorea

ALBA Beamline 20 LOREA

erlab.io.plugins.lorea.LOREALoader

maestro

ALS Beamline 7.0.2.1 MAESTRO

erlab.io.plugins.maestro.MAESTROMicroLoader

mbs

MB Scientific .txt and .krx files

erlab.io.plugins.mbs.MBSLoader

merlin

ALS Beamline 4.0.3 MERLIN

erlab.io.plugins.merlin.MERLINLoader

pal4a1

PAL Beamline 4A1

erlab.io.plugins.pal4a1.PAL4A1Loader

snu1

System 1 at Seoul National University

erlab.io.plugins.snu1.System1Loader

ssrl52

SSRL Beamline 5-2

erlab.io.plugins.ssrl52.SSRL52Loader

tps39a

NSRRC TPS Beamline 39A

erlab.io.plugins.tps39a.TPS39ALoader

Loader names are case-sensitive. Check the registry in your installed environment because the available loaders depend on the ERLabPy version and installed plugins:

import erlab

erlab.io.loaders

Replace the loader name, data directory, and identifier with values for the experiment. The identifier can be a scan number or a file name when the selected loader supports that form:

identifier = 42
with erlab.io.loader_context("merlin", data_dir="/path/to/data"):
    data = erlab.io.load(identifier)

Inspect the result before analysis:

data
data.coords
data.attrs

Confirm that the dimensions, coordinates, and acquisition metadata match the measurement. Before momentum conversion or Fermi edge fitting, also confirm that the result follows the ARPES data conventions.

Keep the context open when you must load several scans from the same experiment:

with erlab.io.loader_context("merlin", data_dir="/path/to/data"):
    data_1 = erlab.io.load(1)
    data_2 = erlab.io.load(2)

If ERLabPy cannot find the scan, confirm the selected loader, the accepted identifier form, and the data directory. If the endstation is not listed, follow the procedure in Implementing a data loader plugin instead of forcing the file through a different loader. See ARPES data conventions for the coordinate and metadata names required by ARPES-specific tools.

Loading data from several experiments in one notebook

Use a separate erlab.io.loader_context() for each experiment. The context limits the selected loader and data directory to its with block:

import erlab

with erlab.io.loader_context("merlin", data_dir="/data/merlin-experiment"):
    merlin_map = erlab.io.load(42)

with erlab.io.loader_context("i05", data_dir="/data/i05-experiment"):
    i05_map = erlab.io.load(17)

Use the same pattern when two experiments use the same loader but have different data directories. Give each result a name that identifies its source. Then inspect the coordinates and attributes before you compare or combine the data:

merlin_map.coords
merlin_map.attrs

i05_map.coords
i05_map.attrs

Do not depend on the order of repeated erlab.io.set_loader() or erlab.io.set_data_dir() calls. Each call changes the active global setting. A loader context keeps each load operation next to the settings that control it.

Loading data exported from Igor Pro

Load a single .ibw wave or a single-wave .itx file as a DataArray:

import xarray as xr

data = xr.load_dataarray("/path/to/wave.ibw")

For a packed experiment (.pxp or .pxt) that contains several folders or waves, open the hierarchy as a DataTree and select the required node:

with xr.open_datatree("/path/to/experiment.pxp") as experiment:
    print(experiment.groups)
    wave_node = experiment["/folder/wave_name"]
    data = wave_node.dataset["wave_name"].load()

Use a group path listed in groups. Then select the wave variable from that node. The call to load reads the selected wave before the file closes.

For an HDF5 file exported by Igor Pro, select the ERLabPy backend explicitly:

data = xr.load_dataset("/path/to/export.h5", engine="erlab-igor")

Inspect the dimensions, coordinates, and attributes after loading. If a complex packed experiment does not load correctly, export the required wave as .ibw and load that file. See erlab.io.igor.IgorBackendEntrypoint for the supported Igor formats and backend behavior.

Loading a scan stored across multiple files

Use this guide when one scan is stored in files such as f_003_S001.pxt, f_003_S002.pxt, and later sequence files. Select the loader and directory for the experiment:

import erlab

loader = erlab.io.loaders["merlin"]
data_dir = "/path/to/data"

Load and concatenate the complete scan with its scan number or with any file in the scan:

scan = loader.load(3, data_dir=data_dir)
scan_from_first_file = loader.load("f_003_S001.pxt", data_dir=data_dir)
scan_from_second_file = loader.load("f_003_S002.pxt", data_dir=data_dir)

Each call returns the same concatenated data when the selected loader supports multi-file scans.

To load only one file, set single=True:

single_file = loader.load("f_003_S001.pxt", data_dir=data_dir, single=True)

To inspect the scan as separate arrays, disable concatenation:

separate_files = loader.load(3, data_dir=data_dir, combine=False)

If loading one sequence file does not find its companions, confirm that the active loader supports multi-file scans and that every file is in the selected data directory. See erlab.io.dataloader.LoaderBase.load() for the loader arguments.

Adding logbook metadata while loading data

Use an Excel log when acquisition settings are not stored in each data file. Identify the column that matches each file and map the remaining columns to ERLabPy coordinates or attributes:

import erlab
from erlab.io.metadata import ExcelMetadataSource

metadata = ExcelMetadataSource(
    "acquisition-log.xlsx",
    sheet_name="Measurements",
    file_name_column="File",
    coordinate_mapping={
        "Photon Energy": "hv",
        "Temperature": "sample_temp",
    },
    attribute_mapping={"Polarization": "polarization"},
    overwrite=False,
)

with erlab.io.loader_context("merlin", data_dir="/path/to/data"):
    data = erlab.io.load(42, metadata=metadata)

Set overwrite=True only when the logbook values must replace scalar values already stored in the file. Confirm the matched row and the resulting coordinates before using the data for momentum conversion or fitting.

If a nonstandard file name cannot be matched to its row, supply the logbook file number explicitly:

with erlab.io.loader_context("merlin", data_dir="/path/to/data"):
    data = erlab.io.load("custom-name.pxt", metadata=metadata, file_number=42)

For a public Google Sheet, use GoogleSheetsMetadataSource with the same mappings. Give anyone with the link view access. Use row_range= when one sheet contains repeated file numbers from several experiments.

See erlab.io.metadata.ExcelMetadataSource and erlab.io.metadata.GoogleSheetsMetadataSource for matching and range syntax.

Preserving file metadata as coordinates

Use a temporary loader extension when metadata needed during concatenation is stored as a file attribute. For one load, promote the attribute through loader_extensions:

import erlab

with erlab.io.loader_context("merlin", data_dir="/path/to/data"):
    data = erlab.io.load(
        1,
        loader_extensions={"coordinate_attrs": ("scan_number",)},
    )

For several loads, apply the extension only inside a context manager:

with erlab.io.loader_context("merlin", data_dir="/path/to/data"):
    with erlab.io.extend_loader(coordinate_attrs=("scan_number",)):
        data_1 = erlab.io.load(1)
        data_2 = erlab.io.load(2)

Inspect the loaded coordinates before analysis. The metadata key must exist in the files handled by the active loader.

See erlab.io.extend_loader() for the extension fields and their accepted values.

Saving analysis data with coordinates and metadata

Save an xarray DataArray in an HDF5-backed NetCDF file when dimensions, coordinates, and attributes must remain available to Python:

data.to_netcdf("analysis-result.h5", engine="h5netcdf")

Open the file later with xarray:

import xarray as xr

restored = xr.load_dataarray("analysis-result.h5", engine="h5netcdf")

Use erlab.io.igor.save_wave() only when the recipient requires an Igor Binary Wave and the DataArray satisfies the format limits:

import erlab

erlab.io.igor.save_wave(data, "analysis-result.ibw")

An Igor Binary Wave supports at most four uniformly sampled dimensions and does not preserve non-dimensional coordinates.