tsgap API
TSGap — Composable Time-Series Missingness Simulation
A library for simulating realistic missingness patterns in time-series data for imputation benchmarking.
Separates two concepts: 1. MECHANISMS (why data is missing): MCAR, MAR, MNAR 2. PATTERNS (how data is missing): pointwise, block, monotone, decay, markov
- class tsgap.MissingnessSimulator(mechanism, missing_rate, seed=None, **config)[source]
Bases:
objectObject-oriented interface for missingness simulation.
- Parameters:
mechanism (str) – Missingness mechanism (“mcar”, “mar”, “mnar”)
missing_rate (float) – Target missing rate
seed (int, optional) – Random seed
**config (dict) – Mechanism-specific configuration
Examples
>>> sim = MissingnessSimulator("mcar", missing_rate=0.15, seed=42) >>> X_missing, mask = sim.generate(X)
- tsgap.simulate_many_rates(X, mechanism, rates, seed=None, **kwargs)[source]
Simulate missingness at multiple rates.
- Parameters:
X (np.ndarray) – Input data
mechanism (str) – Missingness mechanism
rates (list[float]) – List of missing rates to simulate
seed (int, optional) – Base random seed
**kwargs (dict) – Mechanism-specific parameters
- Returns:
Dictionary mapping rate -> (X_missing, mask)
- Return type:
dict
- tsgap.simulate_missingness(X, mechanism, missing_rate, seed=None, pattern='pointwise', **kwargs)[source]
Simulate missingness in time-series data.
This function separates two concepts: the mechanism, which describes why data is missing, and the pattern, which describes how missing values are arranged over time.
- Parameters:
X (np.ndarray) – Input data of shape (T, D) or (N, T, D)
mechanism (str) – Missingness mechanism. Supported values are “mcar” (Missing Completely At Random), “mar” (Missing At Random, dependent on other variables), and “mnar” (Missing Not At Random, dependent on the value itself).
missing_rate (float) – Target fraction of missing values (0.0 to 1.0) Applied to eligible (non-NaN) entries
seed (int, optional) – Random seed for reproducibility
pattern (str, optional) – Missingness pattern. Supported values are “pointwise” (scattered individual points), “block” (contiguous segments), “monotone” (once missing, stays missing), “decay” (missingness increases over time), “markov” (temporally dependent flickering), and “gilbert_elliott” (bursty two-state loss with leaky good/bad periods; MCAR only). Aliases include “point”/”scattered” for pointwise, “contiguous” for block, “dropout” for monotone, “degradation” for decay, “flickering” for markov, and “gilbert-elliott”/”gilbert”/”burst” for gilbert_elliott.
**kwargs (dict) – Mechanism- and pattern-specific parameters. Common options include target, driver_dims, driver_weights, strength, base_rate, direction, mnar_mode, block_len, block_frac, block_density, decay_rate, decay_center, persist, bad_loss, and good_loss. See the user documentation for pattern-specific details and constraints.
- Returns:
X_missing (np.ndarray) – Data with NaNs inserted (same shape as X)
mask (np.ndarray) – Boolean mask (True=observed, False=missing)
- Return type:
tuple[ndarray, ndarray]
Examples
>>> # MCAR with point-wise pattern (default) >>> X_missing, mask = simulate_missingness(X, "mcar", 0.15, seed=42)
>>> # MAR with block pattern (sensor dropout depends on activity) >>> X_missing, mask = simulate_missingness( ... X, "mar", 0.25, seed=42, ... driver_dims=[0], pattern="block", block_len=10 ... )
>>> # Block length can also scale with the time axis >>> X_missing, mask = simulate_missingness( ... X, "mcar", 0.20, seed=42, ... pattern="block", block_frac=0.02 ... )
>>> # MNAR with block pattern (extreme values cause sensor failure) >>> X_missing, mask = simulate_missingness( ... X, "mnar", 0.20, seed=42, ... mnar_mode="extreme", pattern="block" ... )
Core
Core API for missingness simulation.
- class tsgap.core.MissingnessSimulator(mechanism, missing_rate, seed=None, **config)[source]
Bases:
objectObject-oriented interface for missingness simulation.
- Parameters:
mechanism (str) – Missingness mechanism (“mcar”, “mar”, “mnar”)
missing_rate (float) – Target missing rate
seed (int, optional) – Random seed
**config (dict) – Mechanism-specific configuration
Examples
>>> sim = MissingnessSimulator("mcar", missing_rate=0.15, seed=42) >>> X_missing, mask = sim.generate(X)
- tsgap.core.simulate_many_rates(X, mechanism, rates, seed=None, **kwargs)[source]
Simulate missingness at multiple rates.
- Parameters:
X (np.ndarray) – Input data
mechanism (str) – Missingness mechanism
rates (list[float]) – List of missing rates to simulate
seed (int, optional) – Base random seed
**kwargs (dict) – Mechanism-specific parameters
- Returns:
Dictionary mapping rate -> (X_missing, mask)
- Return type:
dict
- tsgap.core.simulate_missingness(X, mechanism, missing_rate, seed=None, pattern='pointwise', **kwargs)[source]
Simulate missingness in time-series data.
This function separates two concepts: the mechanism, which describes why data is missing, and the pattern, which describes how missing values are arranged over time.
- Parameters:
X (np.ndarray) – Input data of shape (T, D) or (N, T, D)
mechanism (str) – Missingness mechanism. Supported values are “mcar” (Missing Completely At Random), “mar” (Missing At Random, dependent on other variables), and “mnar” (Missing Not At Random, dependent on the value itself).
missing_rate (float) – Target fraction of missing values (0.0 to 1.0) Applied to eligible (non-NaN) entries
seed (int, optional) – Random seed for reproducibility
pattern (str, optional) – Missingness pattern. Supported values are “pointwise” (scattered individual points), “block” (contiguous segments), “monotone” (once missing, stays missing), “decay” (missingness increases over time), “markov” (temporally dependent flickering), and “gilbert_elliott” (bursty two-state loss with leaky good/bad periods; MCAR only). Aliases include “point”/”scattered” for pointwise, “contiguous” for block, “dropout” for monotone, “degradation” for decay, “flickering” for markov, and “gilbert-elliott”/”gilbert”/”burst” for gilbert_elliott.
**kwargs (dict) – Mechanism- and pattern-specific parameters. Common options include target, driver_dims, driver_weights, strength, base_rate, direction, mnar_mode, block_len, block_frac, block_density, decay_rate, decay_center, persist, bad_loss, and good_loss. See the user documentation for pattern-specific details and constraints.
- Returns:
X_missing (np.ndarray) – Data with NaNs inserted (same shape as X)
mask (np.ndarray) – Boolean mask (True=observed, False=missing)
- Return type:
tuple[ndarray, ndarray]
Examples
>>> # MCAR with point-wise pattern (default) >>> X_missing, mask = simulate_missingness(X, "mcar", 0.15, seed=42)
>>> # MAR with block pattern (sensor dropout depends on activity) >>> X_missing, mask = simulate_missingness( ... X, "mar", 0.25, seed=42, ... driver_dims=[0], pattern="block", block_len=10 ... )
>>> # Block length can also scale with the time axis >>> X_missing, mask = simulate_missingness( ... X, "mcar", 0.20, seed=42, ... pattern="block", block_frac=0.02 ... )
>>> # MNAR with block pattern (extreme values cause sensor failure) >>> X_missing, mask = simulate_missingness( ... X, "mnar", 0.20, seed=42, ... mnar_mode="extreme", pattern="block" ... )
Mechanisms
Missingness mechanism implementations.
- tsgap.mechanisms.apply_mar(X, missing_rate, existing_nans, driver_dims=None, driver_weights=None, target='all', strength=2.0, base_rate=0.01, direction='positive', rng=None, **kwargs)[source]
Apply MAR (Missing At Random) mechanism.
Missingness depends on driver dimensions (other observed variables).
Masking probability is defined per time step as a logistic function of a driver variable; masking is then sampled independently across eligible features at each time step.
When multiple driver dimensions are specified, the driver signal is computed as a weighted linear combination:
driver_t = Σ_k w_k × X_{t,k}
where w_k are the (normalized) driver_weights. This allows different observed variables to contribute differently to missingness probability. If driver_weights is None, all drivers contribute equally (simple mean).
- Parameters:
X (np.ndarray) – Input data
missing_rate (float) – Target missing rate (applied to eligible entries) Will be clipped to [0.0, 1.0]
existing_nans (np.ndarray) – Boolean mask of existing NaNs (should be np.isnan(X) from original data)
driver_dims (list[int], optional) – Dimensions that drive missingness (default: first dimension)
driver_weights (list[float], optional) – Weights for each driver dimension. Must have same length as driver_dims. Weights are normalized to sum to 1. Default: equal weights (simple mean).
target (str or list[int]) – “all” (default) or list of dimension indices to mask
strength (float) – Dependency strength (higher = stronger dependency, must be >= 0)
base_rate (float) – Minimum probability to avoid all-zeros (should be < missing_rate)
direction (str) – “positive” (high driver -> high missing) or “negative”
rng (np.random.Generator, optional) – Random number generator for reproducibility
- Returns:
mask – Boolean mask (True=observed, False=missing)
- Return type:
np.ndarray
- tsgap.mechanisms.apply_mcar(X, missing_rate, existing_nans, target='all', rng=None, **kwargs)[source]
Apply MCAR (Missing Completely At Random) mechanism.
- Parameters:
X (np.ndarray) – Input data
missing_rate (float) – Target missing rate (applied to eligible non-NaN entries) Will be clipped to [0.0, 1.0]
existing_nans (np.ndarray) – Boolean mask of existing NaNs (should be np.isnan(X) from original data)
target (str or list[int]) – “all” (default) or list of dimension indices to mask
rng (np.random.Generator, optional) – Random number generator for reproducibility
- Returns:
mask – Boolean mask (True=observed, False=missing)
- Return type:
np.ndarray
- tsgap.mechanisms.apply_mnar(X, missing_rate, existing_nans, mnar_mode='extreme', target='all', strength=2.0, rng=None, **kwargs)[source]
Apply MNAR (Missing Not At Random) mechanism.
Missingness depends on the value itself.
- Parameters:
X (np.ndarray) – Input data
missing_rate (float) – Target missing rate (applied to eligible entries) Will be clipped to [0.0, 1.0]
existing_nans (np.ndarray) – Boolean mask of existing NaNs (should be np.isnan(X) from original data)
mnar_mode (str) – “high” (high values missing), “low” (low values missing), or “extreme” (extreme values missing)
target (str or list[int]) – “all” (default) or list of dimension indices to mask
strength (float) – Dependency strength (must be >= 0)
rng (np.random.Generator, optional) – Random number generator for reproducibility
- Returns:
mask – Boolean mask (True=observed, False=missing)
- Return type:
np.ndarray
Patterns
Missing data patterns (HOW data is missing).
- tsgap.patterns.apply_block_pattern(mask, shape, block_len=10, block_frac=None, block_density=1.0, eligible_mask=None, forced_missing=None, rng=None, **kwargs)[source]
Apply block (contiguous) missingness pattern.
Converts some point-wise missingness into contiguous blocks. Simulates realistic sensor dropout periods.
- Parameters:
mask (np.ndarray) – Initial boolean mask (True=observed, False=missing)
shape (tuple) – Shape of the data
block_len (int) – Length of each missing block (in time steps)
block_frac (float or tuple[float, float], optional) – Relative block length as a fraction of the time axis (0.0, 1.0]. If a tuple is provided, a new fraction is sampled uniformly from
(min_frac, max_frac)for each block. If provided, overrides block_len.block_density (float) – Fraction of total missingness allocated to blocks (0.0 to 1.0)
rng (np.random.Generator, optional) – Random number generator for reproducibility
eligible_mask (ndarray | None)
forced_missing (ndarray | None)
- Returns:
mask – Modified mask with block patterns
- Return type:
np.ndarray
- tsgap.patterns.apply_gilbert_elliott_pattern(mask, shape, persist=0.8, bad_loss=1.0, good_loss=0.0, eligible_mask=None, forced_missing=None, rng=None, **kwargs)[source]
Apply the Gilbert-Elliott burst-loss pattern.
A 2-state hidden Markov model widely used to model bursty packet loss in telecommunications. Each (sample, dimension) series alternates between a “good” state (low loss) and a “bad” state (high loss). Unlike the
markovpattern—which is the degenerate case where the bad state is always missing and the good state is never missing—the Gilbert-Elliott model produces ragged bursts: bad periods still let some values through, and good periods can have occasional dropouts.The hidden state evolves as a 2-state Markov chain:
P(bad at t | bad at t-1) = persist P(bad at t | good at t-1) = p_onset
Within each state, a value is missing with a state-dependent probability:
P(missing | bad) = bad_loss (h) P(missing | good) = good_loss (k)
p_onsetis calibrated automatically when the target rate is feasible for the chosen state-loss probabilities. Using the stationary bad-state probabilitypi_bad = p_onset / (p_onset + 1 - persist), the achieved rate ispi_bad * bad_loss + (1 - pi_bad) * good_loss. Solving for the requiredpi_badgiven the target raterho:pi_bad = (rho - good_loss) / (bad_loss - good_loss)
- Parameters:
mask (np.ndarray) – Initial boolean mask from mechanism (True=observed, False=missing).
shape (tuple) – Shape of the data.
persist (float) – Probability of staying in the bad state, range [0, 1). Higher values create longer bursts. Default 0.8.
bad_loss (float) – Probability of a value being missing while in the bad state (h), range (0, 1]. Default 1.0 (bad state always loses).
good_loss (float) – Probability of a value being missing while in the good state (k), range [0, 1). Must be strictly less than
bad_loss. Default 0.0 (good state never loses).rng (np.random.Generator, optional) – Random number generator for reproducibility.
eligible_mask (ndarray | None)
forced_missing (ndarray | None)
- Returns:
mask – Modified mask with Gilbert-Elliott burst-loss structure.
- Return type:
np.ndarray
- tsgap.patterns.apply_markov_pattern(mask, shape, persist=0.8, eligible_mask=None, forced_missing=None, rng=None, **kwargs)[source]
Apply Markov chain temporal dependence pattern.
Missingness at time t depends on whether t-1 was missing, creating realistic “flickering” on/off patterns common in wearable sensor data.
Governed by a 2-state Markov chain per (sample, dimension) series:
P(missing at t | observed at t-1) = p_onset P(missing at t | missing at t-1) = p_persist
The persist parameter controls “stickiness” — how likely a missing state is to continue. Higher values create longer missing bursts; lower values create rapid flickering.
p_onset is automatically calibrated from the target missing count using the stationary distribution:
π_missing = p_onset / (p_onset + 1 - p_persist)
Solving for p_onset:
p_onset = π_missing × (1 - p_persist) / (1 - π_missing)
The chain is simulated independently for each (sample, dimension) series. The mechanism mask’s total missing count is preserved approximately.
- Parameters:
mask (np.ndarray) – Initial boolean mask from mechanism (True=observed, False=missing)
shape (tuple) – Shape of the data
persist (float) – Probability of staying in the missing state once entered. Range [0, 1). Higher = longer missing bursts. Default 0.8 creates moderate-length bursts. Must be strictly less than 1.0.
rng (np.random.Generator, optional) – Random number generator for reproducibility
eligible_mask (ndarray | None)
forced_missing (ndarray | None)
- Returns:
mask – Modified mask with Markov-chain temporal dependence
- Return type:
np.ndarray
- tsgap.patterns.apply_monotone_pattern(mask, shape, eligible_mask=None, forced_missing=None, rng=None, **kwargs)[source]
Apply monotone missingness pattern.
Once a dimension goes missing at time t, it stays missing for all t’ > t. Models participant dropout in longitudinal studies and clinical trials.
The mechanism mask is used to determine how much each dimension should be missing (its missing density). Dimensions that the mechanism targeted more heavily get earlier dropout times. This preserves the mechanism’s influence: under MAR, dimensions driven by high-valued drivers drop out earlier; under MNAR, dimensions with more extreme values drop out earlier; under MCAR, dropout times are roughly uniform.
The total missing count is preserved by distributing the budget across dimensions proportionally to their mechanism-assigned missing densities.
- Parameters:
mask (np.ndarray) – Initial boolean mask from mechanism (True=observed, False=missing)
shape (tuple) – Shape of the data
rng (np.random.Generator, optional) – Random number generator (not used, kept for API consistency)
eligible_mask (ndarray | None)
forced_missing (ndarray | None)
- Returns:
mask – Modified mask with monotone constraint enforced
- Return type:
np.ndarray
- tsgap.patterns.apply_pointwise_pattern(mask, shape=None, rng=None, **kwargs)[source]
Apply point-wise (scattered) missingness pattern.
This is the default pattern - no modification needed. Individual points are missing independently.
- Parameters:
mask (np.ndarray) – Boolean mask (True=observed, False=missing)
shape (tuple, optional) – Shape of the data (not used, for API consistency)
rng (np.random.Generator, optional) – Random number generator (not used, for API consistency)
- Returns:
mask – Unmodified mask (point-wise is the default)
- Return type:
np.ndarray
- tsgap.patterns.apply_temporal_decay_pattern(mask, shape, decay_rate=3.0, decay_center=0.7, eligible_mask=None, forced_missing=None, rng=None, **kwargs)[source]
Apply temporal decay missingness pattern.
Missingness probability increases over time, modeling sensor degradation, battery drain, or participant fatigue. Early timesteps have low missingness; later timesteps have high missingness.
Uses a sigmoid ramp over the time axis:
w(t) = σ(decay_rate × (t_norm - decay_center))
where t_norm ∈ [0, 1] is the normalized time position. The mechanism mask is then resampled with time-weighted probabilities to preserve the overall missing count.
- Parameters:
mask (np.ndarray) – Initial boolean mask from mechanism (True=observed, False=missing)
shape (tuple) – Shape of the data
decay_rate (float) – Steepness of the temporal ramp (higher = sharper transition). Default 3.0 gives a smooth S-curve.
decay_center (float) – Normalized time position (0-1) where missingness reaches 50%. Default 0.7 means most missingness concentrates in the last 30%.
rng (np.random.Generator, optional) – Random number generator for reproducibility
eligible_mask (ndarray | None)
forced_missing (ndarray | None)
- Returns:
mask – Modified mask with temporally increasing missingness
- Return type:
np.ndarray