Database and data sources

vaft.database is the I/O layer of VAFT. It has two independent back-ends, and knowing which one you are talking to explains almost every argument on this page:

Back-end Module What a shot looks like What you get back
HSDS — remote IMAS HDF5 store vaft.database.ods, vaft.database.ids, vaft.database.utils a folder hdf5://{directory}/{shot}/ holding master.h5 plus one <ids_name>.h5 per IDS an OMAS ODS, or a native IMAS IDSToplevel
Raw DAQ — VEST MySQL database vaft.database.raw rows in shotDataWaveform_* addressed by integer field code (time, data) NumPy arrays
flowchart LR
    HSDS[(VEST HSDS<br/>IMAS HDF5 images)]
    SQL[(VEST MySQL<br/>raw DAQ waveforms)]
    subgraph db["vaft.database"]
        ods["ods.load_ods / ods.save_ods"]
        ids["ids.load / ids.save"]
        raw["raw.load_raw"]
    end
    HSDS -->|hsget| ods
    HSDS -->|hsget| ids
    SQL -->|mysql-connector| raw
    ods --> ODS["omas.ODS"]
    ids --> IDS["imas IDSToplevel"]
    raw --> ARR["(time, data) ndarrays"]

The IMAS/ODS data model itself is described in Data structures; this page is about getting the data.

The flat namespace

vaft/__init__.py and vaft/database/__init__.py are both lazy. import vaft is enough — you never need import vaft.database. The database package also resolves unknown attributes by scanning its submodules in the order ods, ids, raw, utils, so helpers defined in raw.py or utils.py are reachable flat:

import vaft

vaft.database.is_connect()          # defined in database/utils.py
vaft.database.exist_shot('public')  # defined in database/utils.py
vaft.database.load_raw(39915, 102)  # defined in database/raw.py

The supported high-level remote API is load, open, and save. The raw, filedb, ods, ids, and utils modules remain available for specialized and compatibility workflows.

Connecting to HSDS

h5pyd is a declared VAFT dependency on develop; a normal pip install . installs the compatible client and its hsconfigure command. Do not install a separately pinned --no-deps copy.

Then write your credentials with hsconfigure:

hsconfigure
Field Value
Server endpoint http://147.46.36.244:5101
Username / Password reader / test (read-only public account)

This writes ~/.hscfg; h5pyd also recognizes a project-local .hscfg. Never commit that file. Check the connection from Python:

import vaft

vaft.database.is_connect()   # True when the HSDS server reports state == "READY"

Local OMAS and IMAS operations do not require HSDS at all; use vaft.omas.load/save or vaft.imas.load/save for those paths. See Start here for the full environment setup.

Listing shots

vaft.database.exist_shot()                       # numeric shot folders in /public/, newest first
vaft.database.exist_shot('public', shot=39915)   # True / False
vaft.database.exist_shot('public', sort=1)       # ascending
vaft.database.exist_shot(data_filter='ts')       # DataFrame of processed Thomson-scattering shots

Full signature:

exist_shot(username=None, shot=None, data_filter=None, sort=-1)
  • username defaults to 'public'. Any other folder name returns that folder’s raw listing.
  • sort is 1 (ascending), -1 (descending, the default) or 0 (unsorted).
  • data_filter='ts' (also 'thomson_scattering') ignores username/shot and returns a pandas.DataFrame with columns Index, Shot Number, Last Processed, Status, read from the processed-shots file hdf5://public_omas/processed_shots.h5 (exposed as vaft.database.utils.PROCESSED_H5_PATH). It returns None when nothing has been processed.
  • Shot folders come back as strings, not ints. A connection failure prints Connection error and returns [] / False.

Eager and lazy HSDS access

vaft.database.load materializes data before returning it. The default representation is an OMAS ODS:

import vaft

ods = vaft.database.load(39915, source="public", paths="magnetics")

time = ods['magnetics.time']
ip = ods['magnetics.ip.0.data']

The current high-level signature is:

load(shot, source="public", *, representation="omas", paths=None,
     occurrence=None, imas_version=None, cache="auto", transport="auto")

Set representation="imas" for native IDS objects. OMAS paths may point at an IDS root or leaf; native IMAS requests accept top-level IDS names only.

# Several shots at once -> list of ODS
ods_list = vaft.database.load([39915, 39916, 39917], paths="magnetics")

# Native IDS materialization
equilibrium = vaft.database.load(
    39915, representation="imas", paths="equilibrium"
)

# Lazy, read-only access: data is fetched only when a path is touched
with vaft.database.open(39915, source="public", paths="equilibrium") as remote:
    times = remote["equilibrium.time"]

open() returns a context-managed lazy OMAS or native-IMAS adapter and never exposes a write operation. Lazy native IMAS access does not convert Data Dictionary versions; use eager load() when conversion is required.

Local OMAS, IMAS, and FileDB paths

Local artifact I/O belongs to the representation modules:

import vaft

ods = vaft.omas.sample_ods()
vaft.omas.save(ods, "/tmp/shot.json.gz")
restored = vaft.omas.load("/tmp/shot.json.gz")

# Native IMAS AL5/URI targets use the corresponding bridge
vaft.imas.save(ods, "imas:hdf5?path=/tmp/imas-entry")
restored = vaft.imas.load("imas:hdf5?path=/tmp/imas-entry")

vaft.database.filedb.FileDB resolves canonical, OMAS-first archive locations without writing:

from vaft.database.filedb import FileDB

db = FileDB.from_config({"filedb": {"root": "/srv/vest.filedb"}})
path = db.path("omas", shot=39915)

Set VAFT_FILEDB_DIR when the configuration uses that environment reference. The audit helpers are read-only and propose legacy-to-canonical mappings without moving data.

Saving to HSDS

save(data, shot, *, target="public", representation=None, occurrence=None,
     imas_version=None, derived_cache="auto")
# Remote writes are admin-restricted; ordinary documentation and analysis are read-only.
uri = vaft.database.save(ods, 39915, target="my-authorized-namespace")

target must be a bare HSDS namespace, never an hdf5:// URI. VAFT infers OMAS versus native IMAS from the supplied object and rejects a conflicting explicit representation.

Native IMAS IDS objects

When you want an IDSToplevel from imas rather than an OMAS ODS, use the IDS pair. Native loading requires the keyword ids_name= — that is what distinguishes it from a directory argument:

# Both forms are equivalent
eq = vaft.database.load(shot=2, ids_name="equilibrium", dd_version="3.41.0")
eq = vaft.database.load_ids(2, "equilibrium", dd_version="3.41.0")

# A list of IDS names returns a dict {name: ids}
idss = vaft.database.load_ids(2, ["equilibrium", "pf_active"])
ids.load(shot, ids_name, directory="public", occurrence=0, dd_version=None, local_dir=None)
ids.save(ids, shot, env="server", path=None, dd_version=None)

ids.load downloads master.h5 and every .h5 it externally links — IMAS-Core refuses to open a data entry with a missing link, even when you ask for a single IDS — then opens the staging directory with imas.DBEntry("imas:hdf5?path=…", "r") and returns dbentry.get(ids_name, occurrence).

Saving a native IDS goes through save_ids, not save:

uri = vaft.database.save_ids(eq, 2, env="server", dd_version="3.41.0")
# -> "hdf5://{username}/2/equilibrium.h5"
  • vaft.database.save / save_ods handle ODS only and have no dd_version parameter; passing one raises TypeError. Use save_ids (or vaft.database.ids.save) for IDS objects.
  • ids.save has no directory argument. The target folder is the logged-in HSDS username (remapped to public when that username is admin), and both {ids_name}.h5 and the regenerated master.h5 are uploaded. The IDS file name comes from the IDS metadata name.
  • env="local" writes to ~/public/imasdb/VEST/3/{shot}/1 by default.

Which load is which

Call Result
vaft.database.load(39915) OMAS ODS
vaft.database.load(39915, "public_omas") OMAS ODS from the public_omas folder
vaft.database.load(2, ids_name="equilibrium") native IMAS IDS
vaft.database.raw.load(39915, 102) (time, data) from the raw SQL database

vaft.database.raw.load is a legacy alias for raw.vest_load living inside raw.py. It is a different function from the package-level vaft.database.load; never write vaft.database.load(shot, field) expecting a waveform.

Raw DAQ signals (MySQL)

Raw signals are addressed by an integer field code and returned as (time, data) NumPy arrays, with time in seconds.

from vaft.database import raw

raw.setup_raw_db()   # interactive: prompts for hostname / username / password
raw.init_pool()      # build the MySQL connection pool

time, ip = raw.load_raw(39915, 102)      # field 102 = Plasma Current

Plasma current of shot #39915

Credentials are stored in ~/.vest/database_raw_info.yaml with the password Fernet-encrypted using ~/.vest/encryption_key.key. setup_raw_db() uses input() prompts, so do not trigger it from an unattended notebook — init_pool() and configuration() will call it if the YAML is missing. load_raw initialises the pool automatically when it has not been built yet.

load_raw(shot, fields=None, max_retries=3, daq_type=None, sample_opt=False)
  • A single int field returns a 1-D data array; a list of fields returns a 2-D array of shape (N, n_fields), column-stacked and truncated to the shortest field.
  • load_raw never raises — it logs and returns None on any failure. Always check the result.
loaded = raw.load_raw(39915, [102, 101, 1])
if loaded is None:
    raise RuntimeError("raw load failed")
time, data = loaded
ip = data[:, 0]     # column order follows the requested field list

Finding field codes and shots

raw.name(102)                             # -> (field name, remark) from the shotDataField table
raw.vest_load_by_name(39915, "Plasma Current")   # load by human name (alias: raw.vest_loadn)
raw.get_all_field_codes_for_shot(39915)   # every field code recorded for the shot
raw.last_shot()                           # highest shot number in the database
raw.date_from_shot(39915)                 # ('YYYY-MM-DD', datetime)
raw.shots_from_date('2023-06-01')         # [shot, shot, ...]
raw.plot(39915, [102, 101])               # matplotlib quick-look

All of these need init_pool() first; they print an error and return None / [] otherwise. vest_load_by_name resolves names through the packaged lookup table vaft/data/legacy/sql_table.txt (a JSON mapping such as {"TF Current": 1, "Plasma Current": 102, …}), also exposed as raw.SQL_TABLE_PATH:

import json
from vaft.database import raw

with open(raw.SQL_TABLE_PATH, "r", encoding="utf-8") as f:
    signal_to_field = json.load(f)

get_all_field_codes_for_shot omits raw.EXCLUDED_FIELD_CODES = {110, 111, 112, 113} (processed triple-probe signals, which sit on a different time base).

Shot-range and timing rules

Which waveform table a shot lives in depends on its number, and load_raw returns None outside these ranges:

Shot range Table
29349 < shot <= 42190 shotDataWaveform_2
shot > 42190 shotDataWaveform_3

Sampling intervals are raw.FAST_DT = 4e-6 s and raw.SLOW_DT = 4e-5 s, classified against raw.SLOW_DT_THRESHOLD = 5e-6 s. A DAQ trigger-delay correction is added to traces that start at the digitiser origin: 0.24 s for shot < 41446, 0.26 s for shots 41446–41451, 0.24 s for 41452–41659, and 0.26 s from 41660 on.

Working offline

load_raw can read gzipped-JSON dumps instead of MySQL, which is how the test suite and the processing pipelines run without database access. Two environment variables control it:

Variable Meaning
VAFT_RAW_SAMPLE_PATH Path template for the archive; {shot} is substituted, e.g. vaft/data/legacy/shot_{shot}.json.gz
VAFT_RAW_OFFLINE_ONLY 1/true/yes/on forbids any live SQL access; load_raw returns None when no archive exists

Resolution order inside load_raw is: an explicit sample_opt string → VAFT_RAW_SAMPLE_PATH → (unless offline-only) the MySQL pool. raw.raw_offline_only() reports the current mode, and raw.sql_loading_available() reports whether the MySQL driver imported at all.

import vaft
from vaft.database import raw

SAMPLE_PATH = vaft.data.data_path("legacy/shot_44740.json.gz")   # packaged archive

loaded = raw.load_raw(44740, 102, sample_opt=str(SAMPLE_PATH))
time, plasma_current = loaded

Produce your own archive for a shot with:

raw.init_pool()
raw.dump_all_raw_signals_for_shot(shot=44740, output_path="vest_raw_44740.json.gz")
dump_all_raw_signals_for_shot(shot, output_path=None, max_retries=3,
                              daq_type=0, slow_dt_threshold=5e-6, plot_opt=False)

It writes {"shot": n, "fields": {code: {"type": "fast" | "slow", "data": [...]}}} as gzipped JSON, defaulting to ./vest_raw_{shot}.json.gz (a .gz suffix is appended if you omit it) and returning a bool. plot_opt=True also saves an overview figure. raw.compare_db_and_dumped_raw_signals_for_shot overlays the database and dumped traces for QA.

The pipeline scripts use exactly this pattern — generate_raw_db_dump.py creates the archive, and generate_diagnostics_ods.py sets VAFT_RAW_SAMPLE_PATH and VAFT_RAW_OFFLINE_ONLY before building the diagnostics ODS.

Where to go next

Notebooks

Source

results matching ""

    No results matching ""