tsgap API

TSGap — Composable Time-Series Missingness Simulation

A library for simulating realistic missingness patterns in time-series data for imputation benchmarking.

Separates two concepts: 1. MECHANISMS (why data is missing): MCAR, MAR, MNAR 2. PATTERNS (how data is missing): pointwise, block, monotone, decay, markov

class tsgap.MissingnessSimulator(mechanism, missing_rate, seed=None, **config)[source]

Bases: object

Object-oriented interface for missingness simulation.

Parameters:
  • mechanism (str) – Missingness mechanism (“mcar”, “mar”, “mnar”)

  • missing_rate (float) – Target missing rate

  • seed (int, optional) – Random seed

  • **config (dict) – Mechanism-specific configuration

Examples

>>> sim = MissingnessSimulator("mcar", missing_rate=0.15, seed=42)
>>> X_missing, mask = sim.generate(X)
generate(X)[source]

Generate missingness for input data.

Parameters:

X (np.ndarray) – Input data

Returns:

  • X_missing (np.ndarray) – Data with missingness

  • mask (np.ndarray) – Boolean mask

Return type:

tuple[ndarray, ndarray]

tsgap.simulate_many_rates(X, mechanism, rates, seed=None, **kwargs)[source]

Simulate missingness at multiple rates.

Parameters:
  • X (np.ndarray) – Input data

  • mechanism (str) – Missingness mechanism

  • rates (list[float]) – List of missing rates to simulate

  • seed (int, optional) – Base random seed

  • **kwargs (dict) – Mechanism-specific parameters

Returns:

Dictionary mapping rate -> (X_missing, mask)

Return type:

dict

tsgap.simulate_missingness(X, mechanism, missing_rate, seed=None, pattern='pointwise', **kwargs)[source]

Simulate missingness in time-series data.

This function separates two concepts: the mechanism, which describes why data is missing, and the pattern, which describes how missing values are arranged over time.

Parameters:
  • X (np.ndarray) – Input data of shape (T, D) or (N, T, D)

  • mechanism (str) – Missingness mechanism. Supported values are “mcar” (Missing Completely At Random), “mar” (Missing At Random, dependent on other variables), and “mnar” (Missing Not At Random, dependent on the value itself).

  • missing_rate (float) – Target fraction of missing values (0.0 to 1.0) Applied to eligible (non-NaN) entries

  • seed (int, optional) – Random seed for reproducibility

  • pattern (str, optional) – Missingness pattern. Supported values are “pointwise” (scattered individual points), “block” (contiguous segments), “monotone” (once missing, stays missing), “decay” (missingness increases over time), “markov” (temporally dependent flickering), and “gilbert_elliott” (bursty two-state loss with leaky good/bad periods; MCAR only). Aliases include “point”/”scattered” for pointwise, “contiguous” for block, “dropout” for monotone, “degradation” for decay, “flickering” for markov, and “gilbert-elliott”/”gilbert”/”burst” for gilbert_elliott.

  • **kwargs (dict) – Mechanism- and pattern-specific parameters. Common options include target, driver_dims, driver_weights, strength, base_rate, direction, mnar_mode, block_len, block_frac, block_density, decay_rate, decay_center, persist, bad_loss, and good_loss. See the user documentation for pattern-specific details and constraints.

Returns:

  • X_missing (np.ndarray) – Data with NaNs inserted (same shape as X)

  • mask (np.ndarray) – Boolean mask (True=observed, False=missing)

Return type:

tuple[ndarray, ndarray]

Examples

>>> # MCAR with point-wise pattern (default)
>>> X_missing, mask = simulate_missingness(X, "mcar", 0.15, seed=42)
>>> # MAR with block pattern (sensor dropout depends on activity)
>>> X_missing, mask = simulate_missingness(
...     X, "mar", 0.25, seed=42,
...     driver_dims=[0], pattern="block", block_len=10
... )
>>> # Block length can also scale with the time axis
>>> X_missing, mask = simulate_missingness(
...     X, "mcar", 0.20, seed=42,
...     pattern="block", block_frac=0.02
... )
>>> # MNAR with block pattern (extreme values cause sensor failure)
>>> X_missing, mask = simulate_missingness(
...     X, "mnar", 0.20, seed=42,
...     mnar_mode="extreme", pattern="block"
... )

Core

Core API for missingness simulation.

class tsgap.core.MissingnessSimulator(mechanism, missing_rate, seed=None, **config)[source]

Bases: object

Object-oriented interface for missingness simulation.

Parameters:
  • mechanism (str) – Missingness mechanism (“mcar”, “mar”, “mnar”)

  • missing_rate (float) – Target missing rate

  • seed (int, optional) – Random seed

  • **config (dict) – Mechanism-specific configuration

Examples

>>> sim = MissingnessSimulator("mcar", missing_rate=0.15, seed=42)
>>> X_missing, mask = sim.generate(X)
generate(X)[source]

Generate missingness for input data.

Parameters:

X (np.ndarray) – Input data

Returns:

  • X_missing (np.ndarray) – Data with missingness

  • mask (np.ndarray) – Boolean mask

Return type:

tuple[ndarray, ndarray]

tsgap.core.simulate_many_rates(X, mechanism, rates, seed=None, **kwargs)[source]

Simulate missingness at multiple rates.

Parameters:
  • X (np.ndarray) – Input data

  • mechanism (str) – Missingness mechanism

  • rates (list[float]) – List of missing rates to simulate

  • seed (int, optional) – Base random seed

  • **kwargs (dict) – Mechanism-specific parameters

Returns:

Dictionary mapping rate -> (X_missing, mask)

Return type:

dict

tsgap.core.simulate_missingness(X, mechanism, missing_rate, seed=None, pattern='pointwise', **kwargs)[source]

Simulate missingness in time-series data.

This function separates two concepts: the mechanism, which describes why data is missing, and the pattern, which describes how missing values are arranged over time.

Parameters:
  • X (np.ndarray) – Input data of shape (T, D) or (N, T, D)

  • mechanism (str) – Missingness mechanism. Supported values are “mcar” (Missing Completely At Random), “mar” (Missing At Random, dependent on other variables), and “mnar” (Missing Not At Random, dependent on the value itself).

  • missing_rate (float) – Target fraction of missing values (0.0 to 1.0) Applied to eligible (non-NaN) entries

  • seed (int, optional) – Random seed for reproducibility

  • pattern (str, optional) – Missingness pattern. Supported values are “pointwise” (scattered individual points), “block” (contiguous segments), “monotone” (once missing, stays missing), “decay” (missingness increases over time), “markov” (temporally dependent flickering), and “gilbert_elliott” (bursty two-state loss with leaky good/bad periods; MCAR only). Aliases include “point”/”scattered” for pointwise, “contiguous” for block, “dropout” for monotone, “degradation” for decay, “flickering” for markov, and “gilbert-elliott”/”gilbert”/”burst” for gilbert_elliott.

  • **kwargs (dict) – Mechanism- and pattern-specific parameters. Common options include target, driver_dims, driver_weights, strength, base_rate, direction, mnar_mode, block_len, block_frac, block_density, decay_rate, decay_center, persist, bad_loss, and good_loss. See the user documentation for pattern-specific details and constraints.

Returns:

  • X_missing (np.ndarray) – Data with NaNs inserted (same shape as X)

  • mask (np.ndarray) – Boolean mask (True=observed, False=missing)

Return type:

tuple[ndarray, ndarray]

Examples

>>> # MCAR with point-wise pattern (default)
>>> X_missing, mask = simulate_missingness(X, "mcar", 0.15, seed=42)
>>> # MAR with block pattern (sensor dropout depends on activity)
>>> X_missing, mask = simulate_missingness(
...     X, "mar", 0.25, seed=42,
...     driver_dims=[0], pattern="block", block_len=10
... )
>>> # Block length can also scale with the time axis
>>> X_missing, mask = simulate_missingness(
...     X, "mcar", 0.20, seed=42,
...     pattern="block", block_frac=0.02
... )
>>> # MNAR with block pattern (extreme values cause sensor failure)
>>> X_missing, mask = simulate_missingness(
...     X, "mnar", 0.20, seed=42,
...     mnar_mode="extreme", pattern="block"
... )

Mechanisms

Missingness mechanism implementations.

tsgap.mechanisms.apply_mar(X, missing_rate, existing_nans, driver_dims=None, driver_weights=None, target='all', strength=2.0, base_rate=0.01, direction='positive', rng=None, **kwargs)[source]

Apply MAR (Missing At Random) mechanism.

Missingness depends on driver dimensions (other observed variables).

Masking probability is defined per time step as a logistic function of a driver variable; masking is then sampled independently across eligible features at each time step.

When multiple driver dimensions are specified, the driver signal is computed as a weighted linear combination:

driver_t = Σ_k w_k × X_{t,k}

where w_k are the (normalized) driver_weights. This allows different observed variables to contribute differently to missingness probability. If driver_weights is None, all drivers contribute equally (simple mean).

Parameters:
  • X (np.ndarray) – Input data

  • missing_rate (float) – Target missing rate (applied to eligible entries) Will be clipped to [0.0, 1.0]

  • existing_nans (np.ndarray) – Boolean mask of existing NaNs (should be np.isnan(X) from original data)

  • driver_dims (list[int], optional) – Dimensions that drive missingness (default: first dimension)

  • driver_weights (list[float], optional) – Weights for each driver dimension. Must have same length as driver_dims. Weights are normalized to sum to 1. Default: equal weights (simple mean).

  • target (str or list[int]) – “all” (default) or list of dimension indices to mask

  • strength (float) – Dependency strength (higher = stronger dependency, must be >= 0)

  • base_rate (float) – Minimum probability to avoid all-zeros (should be < missing_rate)

  • direction (str) – “positive” (high driver -> high missing) or “negative”

  • rng (np.random.Generator, optional) – Random number generator for reproducibility

Returns:

mask – Boolean mask (True=observed, False=missing)

Return type:

np.ndarray

tsgap.mechanisms.apply_mcar(X, missing_rate, existing_nans, target='all', rng=None, **kwargs)[source]

Apply MCAR (Missing Completely At Random) mechanism.

Parameters:
  • X (np.ndarray) – Input data

  • missing_rate (float) – Target missing rate (applied to eligible non-NaN entries) Will be clipped to [0.0, 1.0]

  • existing_nans (np.ndarray) – Boolean mask of existing NaNs (should be np.isnan(X) from original data)

  • target (str or list[int]) – “all” (default) or list of dimension indices to mask

  • rng (np.random.Generator, optional) – Random number generator for reproducibility

Returns:

mask – Boolean mask (True=observed, False=missing)

Return type:

np.ndarray

tsgap.mechanisms.apply_mnar(X, missing_rate, existing_nans, mnar_mode='extreme', target='all', strength=2.0, rng=None, **kwargs)[source]

Apply MNAR (Missing Not At Random) mechanism.

Missingness depends on the value itself.

Parameters:
  • X (np.ndarray) – Input data

  • missing_rate (float) – Target missing rate (applied to eligible entries) Will be clipped to [0.0, 1.0]

  • existing_nans (np.ndarray) – Boolean mask of existing NaNs (should be np.isnan(X) from original data)

  • mnar_mode (str) – “high” (high values missing), “low” (low values missing), or “extreme” (extreme values missing)

  • target (str or list[int]) – “all” (default) or list of dimension indices to mask

  • strength (float) – Dependency strength (must be >= 0)

  • rng (np.random.Generator, optional) – Random number generator for reproducibility

Returns:

mask – Boolean mask (True=observed, False=missing)

Return type:

np.ndarray

Patterns

Missing data patterns (HOW data is missing).

tsgap.patterns.apply_block_pattern(mask, shape, block_len=10, block_frac=None, block_density=1.0, eligible_mask=None, forced_missing=None, rng=None, **kwargs)[source]

Apply block (contiguous) missingness pattern.

Converts some point-wise missingness into contiguous blocks. Simulates realistic sensor dropout periods.

Parameters:
  • mask (np.ndarray) – Initial boolean mask (True=observed, False=missing)

  • shape (tuple) – Shape of the data

  • block_len (int) – Length of each missing block (in time steps)

  • block_frac (float or tuple[float, float], optional) – Relative block length as a fraction of the time axis (0.0, 1.0]. If a tuple is provided, a new fraction is sampled uniformly from (min_frac, max_frac) for each block. If provided, overrides block_len.

  • block_density (float) – Fraction of total missingness allocated to blocks (0.0 to 1.0)

  • rng (np.random.Generator, optional) – Random number generator for reproducibility

  • eligible_mask (ndarray | None)

  • forced_missing (ndarray | None)

Returns:

mask – Modified mask with block patterns

Return type:

np.ndarray

tsgap.patterns.apply_gilbert_elliott_pattern(mask, shape, persist=0.8, bad_loss=1.0, good_loss=0.0, eligible_mask=None, forced_missing=None, rng=None, **kwargs)[source]

Apply the Gilbert-Elliott burst-loss pattern.

A 2-state hidden Markov model widely used to model bursty packet loss in telecommunications. Each (sample, dimension) series alternates between a “good” state (low loss) and a “bad” state (high loss). Unlike the markov pattern—which is the degenerate case where the bad state is always missing and the good state is never missing—the Gilbert-Elliott model produces ragged bursts: bad periods still let some values through, and good periods can have occasional dropouts.

The hidden state evolves as a 2-state Markov chain:

P(bad at t | bad at t-1) = persist P(bad at t | good at t-1) = p_onset

Within each state, a value is missing with a state-dependent probability:

P(missing | bad) = bad_loss (h) P(missing | good) = good_loss (k)

p_onset is calibrated automatically when the target rate is feasible for the chosen state-loss probabilities. Using the stationary bad-state probability pi_bad = p_onset / (p_onset + 1 - persist), the achieved rate is pi_bad * bad_loss + (1 - pi_bad) * good_loss. Solving for the required pi_bad given the target rate rho:

pi_bad = (rho - good_loss) / (bad_loss - good_loss)

Parameters:
  • mask (np.ndarray) – Initial boolean mask from mechanism (True=observed, False=missing).

  • shape (tuple) – Shape of the data.

  • persist (float) – Probability of staying in the bad state, range [0, 1). Higher values create longer bursts. Default 0.8.

  • bad_loss (float) – Probability of a value being missing while in the bad state (h), range (0, 1]. Default 1.0 (bad state always loses).

  • good_loss (float) – Probability of a value being missing while in the good state (k), range [0, 1). Must be strictly less than bad_loss. Default 0.0 (good state never loses).

  • rng (np.random.Generator, optional) – Random number generator for reproducibility.

  • eligible_mask (ndarray | None)

  • forced_missing (ndarray | None)

Returns:

mask – Modified mask with Gilbert-Elliott burst-loss structure.

Return type:

np.ndarray

tsgap.patterns.apply_markov_pattern(mask, shape, persist=0.8, eligible_mask=None, forced_missing=None, rng=None, **kwargs)[source]

Apply Markov chain temporal dependence pattern.

Missingness at time t depends on whether t-1 was missing, creating realistic “flickering” on/off patterns common in wearable sensor data.

Governed by a 2-state Markov chain per (sample, dimension) series:

P(missing at t | observed at t-1) = p_onset P(missing at t | missing at t-1) = p_persist

The persist parameter controls “stickiness” — how likely a missing state is to continue. Higher values create longer missing bursts; lower values create rapid flickering.

p_onset is automatically calibrated from the target missing count using the stationary distribution:

π_missing = p_onset / (p_onset + 1 - p_persist)

Solving for p_onset:

p_onset = π_missing × (1 - p_persist) / (1 - π_missing)

The chain is simulated independently for each (sample, dimension) series. The mechanism mask’s total missing count is preserved approximately.

Parameters:
  • mask (np.ndarray) – Initial boolean mask from mechanism (True=observed, False=missing)

  • shape (tuple) – Shape of the data

  • persist (float) – Probability of staying in the missing state once entered. Range [0, 1). Higher = longer missing bursts. Default 0.8 creates moderate-length bursts. Must be strictly less than 1.0.

  • rng (np.random.Generator, optional) – Random number generator for reproducibility

  • eligible_mask (ndarray | None)

  • forced_missing (ndarray | None)

Returns:

mask – Modified mask with Markov-chain temporal dependence

Return type:

np.ndarray

tsgap.patterns.apply_monotone_pattern(mask, shape, eligible_mask=None, forced_missing=None, rng=None, **kwargs)[source]

Apply monotone missingness pattern.

Once a dimension goes missing at time t, it stays missing for all t’ > t. Models participant dropout in longitudinal studies and clinical trials.

The mechanism mask is used to determine how much each dimension should be missing (its missing density). Dimensions that the mechanism targeted more heavily get earlier dropout times. This preserves the mechanism’s influence: under MAR, dimensions driven by high-valued drivers drop out earlier; under MNAR, dimensions with more extreme values drop out earlier; under MCAR, dropout times are roughly uniform.

The total missing count is preserved by distributing the budget across dimensions proportionally to their mechanism-assigned missing densities.

Parameters:
  • mask (np.ndarray) – Initial boolean mask from mechanism (True=observed, False=missing)

  • shape (tuple) – Shape of the data

  • rng (np.random.Generator, optional) – Random number generator (not used, kept for API consistency)

  • eligible_mask (ndarray | None)

  • forced_missing (ndarray | None)

Returns:

mask – Modified mask with monotone constraint enforced

Return type:

np.ndarray

tsgap.patterns.apply_pointwise_pattern(mask, shape=None, rng=None, **kwargs)[source]

Apply point-wise (scattered) missingness pattern.

This is the default pattern - no modification needed. Individual points are missing independently.

Parameters:
  • mask (np.ndarray) – Boolean mask (True=observed, False=missing)

  • shape (tuple, optional) – Shape of the data (not used, for API consistency)

  • rng (np.random.Generator, optional) – Random number generator (not used, for API consistency)

Returns:

mask – Unmodified mask (point-wise is the default)

Return type:

np.ndarray

tsgap.patterns.apply_temporal_decay_pattern(mask, shape, decay_rate=3.0, decay_center=0.7, eligible_mask=None, forced_missing=None, rng=None, **kwargs)[source]

Apply temporal decay missingness pattern.

Missingness probability increases over time, modeling sensor degradation, battery drain, or participant fatigue. Early timesteps have low missingness; later timesteps have high missingness.

Uses a sigmoid ramp over the time axis:

w(t) = σ(decay_rate × (t_norm - decay_center))

where t_norm ∈ [0, 1] is the normalized time position. The mechanism mask is then resampled with time-weighted probabilities to preserve the overall missing count.

Parameters:
  • mask (np.ndarray) – Initial boolean mask from mechanism (True=observed, False=missing)

  • shape (tuple) – Shape of the data

  • decay_rate (float) – Steepness of the temporal ramp (higher = sharper transition). Default 3.0 gives a smooth S-curve.

  • decay_center (float) – Normalized time position (0-1) where missingness reaches 50%. Default 0.7 means most missingness concentrates in the last 30%.

  • rng (np.random.Generator, optional) – Random number generator for reproducibility

  • eligible_mask (ndarray | None)

  • forced_missing (ndarray | None)

Returns:

mask – Modified mask with temporally increasing missingness

Return type:

np.ndarray