Mechanisms
Installation | Core concepts | Mechanisms | Patterns | Mathematical details | Benchmarking workflow | API reference
Mechanisms describe the relationship between the data and the probability that a value is missing.
For full formulas, normalization rules, and sampling details, see Mathematical details.
MCAR
mcar means Missing Completely At Random. Every eligible position has the same
chance of being masked, independent of values in the array.
X_missing, mask = simulate_missingness(X, "mcar", 0.15, seed=42)
MCAR samples without replacement, so the achieved missing count is exact up to rounding.
Parameters:
Parameter |
Default |
Description |
|---|---|---|
|
|
Dimensions to mask, or |
MAR
mar means Missing At Random. Missingness depends on observed driver
dimensions, not on the value being masked.
X_missing, mask = simulate_missingness(
X, "mar", 0.25, seed=42,
driver_dims=[0], strength=2.0
)
With multiple driver dimensions, driver_weights controls their relative
contribution. Weights are normalized automatically.
X_missing, mask = simulate_missingness(
X, "mar", 0.25, seed=42,
driver_dims=[0, 1],
driver_weights=[0.8, 0.2],
strength=2.0
)
Parameters:
Parameter |
Default |
Description |
|---|---|---|
|
|
Dimensions that drive missingness |
|
|
Optional non-negative weights for drivers |
|
|
Dimensions to mask, or |
|
|
Dependency strength, must be non-negative |
|
|
Minimum probability floor |
|
|
|
MNAR
mnar means Missing Not At Random. Missingness depends on the value itself.
X_missing, mask = simulate_missingness(
X, "mnar", 0.20, seed=42,
mnar_mode="extreme", strength=3.0
)
Modes:
Mode |
Effect |
|---|---|
|
High values are more likely to be missing |
|
Low values are more likely to be missing |
|
Values far from the mean are more likely to be missing |
Parameters:
Parameter |
Default |
Description |
|---|---|---|
|
|
|
|
|
Dimensions to mask, or |
|
|
Dependency strength, must be non-negative |
Rate Calibration
MAR and MNAR use a logistic probability model and calibrate an offset by binary search so the expected missing rate over eligible positions matches the target rate. Because they sample Bernoulli outcomes, achieved rates are approximate and vary more on small arrays.