Skip to content

Transformations

Reusable transforms over harmonized tables. The derived tables (food_expenditures, food_prices, food_quantities, household_characteristics) are built from these at API time, but the module is also a library of standalone derived measures — asset indices, dependency ratios, tropical livestock units, anthropometric z-scores, the World Bank median-price valuation ladder — that you can call directly.

Check here before writing a new derived measure

Several of these have no call site inside the library yet, so searching for a usage will not find them. make api-audit lists exactly which.

transformations

A collection of mappings to transform dataframes.

NutrientCoverageWarning

Bases: UserWarning

A fertilizer input label resolved to NO nitrogen share.

:func:nitrogen_kg keys its nutrient share off the plot_inputs.input label, and a label the map does not recognise contributes NOTHING to the total -- silently, and with the right shape: the frame comes back non-empty, finite and non-negative, so every naive assertion passes. That is the "right shape, no content" failure CLAUDE.md devotes a section to (null_read_audit), and it happened: before the label normalisation landed, Malawi's four canonical fertilizer labels (NPK Fertilizer, Urea Fertilizer, CAN Fertilizer, DAP Fertilizer -- 53,210 rows) matched none of the map's keys, the ONLY label that did match was Other Fertilizer at a nominal share of 0.0, and nitrogen_kg(Malawi) returned 331 rows of exactly 0.0.

Fires for an unmatched label that names itself a fertilizer input -- the corpus's own vocabulary, :data:_FERTILIZER_LABEL_HINTS (fertilizer / fertiliser / manure / compost, plus the Portuguese adubo and French engrais) -- and NOT for Seed / Pesticide / Herbicide, which are correctly outside a nitrogen map and are the bulk of the rows (Malawi Seed alone is 112,296). A warning nobody reads is how the original defect survived, so the firehose form is deliberately not used. Escalates to a louder message when the whole returned column is empty or identically zero.

The full per-label tally always rides on result.attrs['nitrogen_input_match'] whether or not this warns -- that is the coverage check, and it is the twin of attrs['kg_factor_sources']['shipped_matched'] in :func:harvest_kg_factors: a JOIN diagnostic, so "did my labels key correctly?" has an answer that does not depend on a warning firing.

PlotGrainMismatchWarning

Bases: UserWarning

Two plot-level features share no land key, so the join produced nothing.

The corpus runs two plot vocabularies (see :data:_PLOT_LEVELS and :func:_parcel_from_crop_plot), and a transform that joins across them on the wrong grain returns an EMPTY frame rather than an error. Measured: fertilizer_rate(Uganda, on='plot') returned (0, 1) silently, because the ag modules key plots {hhid}-{parcel}-{plot} while plot_features keys them {parcel}_{suffix}; the same call with on='parcel' returns 509 rows.

ShippedFactorWarning

Bases: UserWarning

A shipped_factors table did not join the way its author meant.

Two situations, both of which used to be SILENT and both of which look exactly like a working table from the counts alone:

  • the table matched zero rows -- almost always a vocabulary mismatch on a key (j as a raw WB crop code against a decoded label, a numeric 74.0 against '74', t as an int against '2019-20');
  • the table is keyed on a level crop_production does not carry, so that level is DROPPED and the factor is applied across it -- a one-region table going national, a one-condition table serving every condition. The multi-row form of this already raised (the surviving keys become ambiguous); the single-row form did not, so the more careless loader got the weaker signal. This makes both loud.

Its own class so it can be filtered or promoted to an error (warnings.simplefilter('error', ShippedFactorWarning)) independently of the library's other warnings.

StatedVsInferredKgWarning

Bases: UserWarning

The price-ratio inference disagrees with a unit label's OWN stated size.

A label that names its metric content (Sac moyen (50 kg), Gramme, Centilitre) is served its STATED factor -- the label is survey information, the inference is an estimate. When the two disagree by more than :data:STATED_VS_INFERRED_TOLERANCE, the disagreement is evidence about the unit vocabulary (a broken kg baseline, a label that means two sizes) and was previously thrown away. This warning names both numbers.

Its own class so it can be filtered or promoted (warnings.simplefilter('error', StatedVsInferredKgWarning)) independently of the library's other warnings.

UnitLabelCollisionWarning

Bases: UserWarning

Two u spellings that differ only in case got DIFFERENT kg factors.

conversion_to_kgs groups by the raw u label, so Calebasse and calebasse are two units; _get_kg_factors then lower-cases and keeps whichever it merges first. The winner is deterministic (the inference returns a sorted mapping) but ARBITRARY, and the two factors can differ several-fold.

Measured over the 19 cached food_acquired tables (GH #770 review): 16 collision groups in 3 countries, and in every one the two factors DIFFER. Five are INERT -- three kg and two litre, keys that KNOWN_METRIC seeds, so neither inferred value is ever used -- leaving 11 live groups in 3 countries: Burkina Faso 7, Malawi 2, Mali 2. Burkina Faso's Calebasse 0.35 vs calebasse 1.432 is 4.1x; Mali's Unité 1.2 vs unité 0.286 is 4.2x. Only the live groups warn.

Panama is NOT in that list, and was the case that surfaced this: 'Value' -> 0.3489 and 'value' -> 2.3405 was a 6.7x iteration-order accident. Both spellings are currency-denominated, so GH #770 removes them from the inference and the collision with them -- which is why the corpus census above is taken AFTER that fix and still finds 11.

This warns rather than merging, because merging would CHANGE the served factors for those countries -- a data change, not a diagnostic. The fix belongs in the country's own unit vocabulary: canonicalise the spelling where the label is minted, exactly as GH #770 did for currency labels.

UnpriceableRowsWarning

Bases: UserWarning

Emitted when food_prices drops a large share of priceable rows.

Subclasses UserWarning so it can be filtered / promoted to an error (warnings.simplefilter('error', UnpriceableRowsWarning)) independently of the library's other warnings.

ValuationLadderWarning

Bases: UserWarning

The own-production valuation ladder ran on fewer geographic rungs than the canonical v -> District -> Region -> national one.

Emitted by the Country.food_expenditures(valuation=...) path when the country's cluster_features does not carry one of the coarser rungs (or the frame carries no v at all), naming which rungs were available. A shorter ladder is not an error -- median_price_valuation's national fallback is unconditional, so every row still gets a price -- but it means the imputed price is drawn from a coarser market than the caller may assume, so it is said out loud rather than inferred from silence.

age_intervals

age_intervals(age, age_cuts=(4, 9, 14, 19, 31, 51))

Bucket ages into half-open intervals for household_characteristics.

age_cuts is a strictly increasing tuple of interior breakpoints, each strictly positive. Breakpoints may be fractional (e.g. 0.5 to separate neonates from older infants). The tuple partitions ages into len(age_cuts) + 1 half-open buckets:

[0, c_0), [c_0, c_1), ..., [c_{n-1}, inf)

The default (4, 9, 14, 19, 31, 51) reproduces the demographic buckets 00-03, 04-08, ..., 51+ used throughout the library.

Back-compat: earlier releases took (0, 4, 9, ...) with a leading zero. A leading zero is now stripped with a :class:DeprecationWarning so legacy callers keep producing identical buckets.

Parameters:

Name Type Description Default
age array - like

Numeric ages. Negative values fall outside the first bucket and become NaN.

required
age_cuts tuple of positive numbers, strictly increasing

Interior breakpoints between buckets.

(4, 9, 14, 19, 31, 51)

Returns:

Type Description
Categorical

Half-open [a, b) intervals, one per input age.

anthropometry_zscores

anthropometry_zscores(anthropometry, who_reference, *, weight_col='Weight', height_col='Height', age_col='Age_months', sex_col='Sex', male_value='M', female_value='F', flag=True)

WHO-2006 child anthropometric z-scores (WB haz06/waz06/whz06/bmiz06).

METHODOLOGY transform (GAP 5). Reproduces the Stata zscore06 / zanthro call the WB applies (e.g. ETH_ESS1.do:1217-1235) using the WHO 2006 Child Growth Standards: from a child's weight, height/length, age and sex it computes height-for-age (HAZ), weight-for-age (WAZ), weight-for-height (WHZ) and BMI-for-age (BMIZ) z-scores, plus the derived stunting (HAZ < −2) and wasting (WHZ < −2) indicators.

Method (the choice being encoded — documented for review)

The WHO standards distribute each measure as a Box-Cox / LMS family indexed by sex and age (or, for weight-for-height, by sex and height). Given the reference parameters L (Box-Cox power), M (median) and S (coefficient of variation) for a child's (sex, age) cell, the z-score of a measurement y is

z = ((y / M) ** L − 1) / (L · S)            (L ≠ 0)
z = ln(y / M) / S                            (L = 0)

For the WHO 2006 standards, when |z| > 3 the score is recomputed off the SD at ±3 to tame the heavy Box-Cox tail (the WHO "restricted" formula that zscore06 implements):

z = 3 + (y − C₃) / (C₃ − C₂)   for z > 3,   where Cₖ = M·(1+L·S·k)**(1/L)
z = −3 + (y − C₋₃) / (C₋₂ − C₋₃)   for z < −3.

This is the WHO-recommended construction and is exactly what the Stata zscore06 macro does; encoding that method (rather than a plain normal-approximation z) is the choice being documented here. With flag=True (default) z-scores outside the WHO biological-plausibility bounds (:data:_WHO_FLAG_BOUNDS) are set missing, matching the WB's flagged-out implausible measurements.

Reference data (NOT vendored)

The WHO LMS reference tables are not shipped with the library (they are a multi-file WHO dataset under WHO's own terms). The caller supplies them via who_reference — see that parameter. This keeps the transform additive and honest: it encodes the WHO-2006 method and leaves the reference lookup table as an explicit, swappable input, exactly as the Stata macro relies on the externally-installed WHO ado reference.

Parameters:

Name Type Description Default
anthropometry DataFrame

The anthropometry item feature at grain (t, i, v, pid) with Weight (kg) and Height (cm) columns. It must ALSO carry the child's age in months and sex; where the feature itself lacks them (e.g. Tanzania anthropometry has no Age_months), the analyst joins Age_months and Sex from household_roster first — the WB code likewise merges the roster for age/female before the zscore06 call.

required
who_reference dict[str, DataFrame] or callable

The WHO 2006 LMS parameter tables. Either:

  • a dict mapping indicator → DataFrame with columns ['sex', 'x', 'L', 'M', 'S'] where x is age-in-months for 'haz'/'waz'/'bmiz' and height-in-cm for 'whz', and sex is 1/2 (male/female, WHO convention) — keys 'haz', 'waz', 'whz', 'bmiz' (indicators with no table provided are skipped); or
  • a callable who_reference(indicator, sex, x) -> (L, M, S) for callers who back the lookup with pygrowup or the WHO igrowup tables directly.

The igrowup tables are obtainable from https://www.who.int/tools/child-growth-standards/software (the same reference the Stata zscore06 macro installs).

required
weight_col str

Measurement columns (kg, cm).

'Weight'
height_col str

Measurement columns (kg, cm).

'Weight'
age_col str

Child age in months.

'Age_months'
sex_col str

Child sex column, values male_value / female_value.

'Sex'
male_value default 'M' / 'F'

Sex labels in sex_col mapped to WHO 1 / 2.

'M'
female_value default 'M' / 'F'

Sex labels in sex_col mapped to WHO 1 / 2.

'M'
flag bool

Apply the WHO biological-plausibility flagging (set implausible z-scores missing).

True

Returns:

Type Description
DataFrame

anthropometry's index plus columns haz06, waz06, whz06, bmiz06 (whichever the reference supplies) and the derived booleans stunting (haz06 < -2) and wasting (whz06 < -2). Rows outside the WHO age domain (typically 0-60 months) get missing z-scores.

Notes

Divergence from WB: identical method to zscore06; numeric agreement is to the precision of the supplied reference table and the linear interpolation between its tabulated x points (WHO tables are at integer months / 0.1 cm — we interpolate linearly in x, as the WHO macro does). Because the reference is an explicit input, results match the WB exactly when fed the same WHO igrowup tables. bmiz06 uses BMI = weight / (height/100) ** 2.

asset_index

asset_index(assets, split='hh', *, value_col=None, ag_items=None, n_components=1)

First-principal-component asset index (WB ag/hh_asset_index).

METHODOLOGY transform (GAP 8). Reproduces the World Bank factor d_*, pcf + predict construct (e.g. ETH_ESS1.do:1077-1102): reshape the long item-level assets ownership into a household × asset-type ownership matrix, extract its first principal component, and return each household's score on that component as a single wealth index.

Method (the choice being encoded — documented for review)

Stata's factor ..., pcf is principal-component factoring: the factor loadings are the eigenvectors of the correlation matrix, and predict returns the standardised first-component score. We reproduce that exactly with scikit-learn:

  1. Build a household × asset-type binary matrix D of ownership dummies (1 if the HH owns ≥1 of that asset type, else 0). An asset type the HH never reports is a structural 0, matching the WB reshape wide + implicit-zero behaviour.
  2. Drop asset-type columns with no variation (all-0 or all-1) — a constant column has zero correlation contribution and Stata's factor silently ignores it.
  3. Standardise each column to mean 0 / unit variance, then run PCA; operating on standardised columns makes PCA factor the correlation matrix, which is what pcf does.
  4. The index is the projection onto the first principal component (predict after factor). Its sign is arbitrary (PCA sign is not identified); we orient it so the component correlates non-negatively with the number of assets owned — a higher index means more assets, the conventional reading.

The score is per (t, i) and computed within each wave t separately (the WB factors each survey round on its own data), so scores are not comparable in level across waves — only ranks within a wave are.

Parameters:

Name Type Description Default
assets DataFrame

The assets item feature at grain (t, i, j) where j is the asset type. Ownership is inferred from whichever signal is present: a count/Quantity column (> 0 ⇒ owned), else value_col / Value (> 0 or non-missing ⇒ owned), else mere row presence.

required
split (hh, ag)

Which index the WB builds.

  • 'ag' → ag_asset_index: restrict to agricultural asset types (ag_items); the WB keeps only farm-equipment item codes.
  • 'hh' → hh_asset_index: the household (non-ag) asset types — every asset type NOT in ag_items.

When ag_items is None the split is a no-op (the index is built over all asset types) and a single index is returned; pass the country's ag asset-type labels to actually split.

'hh'
value_col str

Column to read ownership from when no count column is present. Defaults to trying 'Quantity' then 'Value'.

None
ag_items collection of str

Asset-type labels (j values) that count as agricultural. Used to partition for split. Country-specific (the WB hard-codes item codes per survey); supply the country's mapping.

None
n_components int

Number of leading components to return (1 = the WB single index). >1 returns Asset_index_1..n for diagnostics.

1

Returns:

Type Description
DataFrame

One Asset_index column (or Asset_index_1..n) indexed by (t, i). Standardised within wave (mean ≈ 0, sd ≈ 1), matching the WB score whose std is ≈ 1.

Notes

Divergence from WB: scikit-learn's eigen-decomposition and Stata's pcf agree on the component direction up to sign and numerical precision, so the index correlates ~1 with the WB column in rank, but absolute values differ by the arbitrary sign (we fix it via the asset-count orientation above) and by tiny scaling/standardisation conventions. Compare via Spearman rank, not equality. Households absent from the assets roster (own nothing of the relevant class) are absent from the result; reindex against the HH universe and fill with the wave minimum if you need them.

conversion_to_kgs

conversion_to_kgs(df, price=['Expenditure'], quantity='Quantity', index=['t', 'm', 'i'], unit_col='u', *, item_col=None, min_reports=FOOD_KG_MIN_REPORTS, min_baseline=FOOD_KG_MIN_BASELINE, min_baseline_tight=FOOD_KG_MIN_BASELINE_TIGHT, tight_tolerance=FOOD_KG_TIGHT_TOLERANCE, baseline_max_spread=FOOD_KG_BASELINE_MAX_SPREAD, _detail=False)

Infer local-unit -> kg conversion factors from price ratios.

For each unit that does not appear in :data:KNOWN_METRIC, this function computes a factor by assuming the price per kilogram should be roughly constant across units for the same item/market. That is: if a "bunch" of item j trades at roughly 2x the unit value of a kg of item j, the inferred factor is 2 kg per bunch.

The mechanics, in three steps:

  1. the baseline -- each row's price per kilogram (price / Kgs) is medianed over a group, giving a reference price for a kilogram;
  2. step 2 -- each row's price per UNIT (price / quantity) is medianed over the same group plus u, over rows whose kilograms are not already known;
  3. the factor -- their ratio, medianed over the remaining axes.

Two things about that chain were wrong before GH #850, and both are fixed here.

Step 2 used to median price itself, not price / quantity, while this docstring described the result as a per-unit price. The factor it returned was therefore kilograms per transaction ROW, equal to kilograms per unit only where Quantity == 1 -- between 4.6% and 41.8% of the corpus's rows. The self-test is the cheap way to see it: asked what a kilogram weighs, the old chain answered between 0.87 and 2.0 depending on the country, and the deviation tracked the median quantity on the rows it looked at.

No step carried the ITEM, so one factor per unit label served a whole country: Malawi's Piece is shared by 133 food items, and the 72 the data can estimate separately run from 1 g to 1.8 kg. Pass item_col='j' for the item-keyed estimate; the default None reproduces the old keys (and the old return type), with the corrected arithmetic.

Labels in :data:_CURRENCY_DENOMINATED_UNITS (e.g. u='Value') are dropped from the whole computation before anything is inferred, so they never appear as a key of the result: their factor is 1/price at (j, t, m) and a per-label constant cannot represent it (GH #770).

The item arm's grouping, and why it is not index + ['j']

The baseline is keyed (t, item) -- the WAVE levels of index plus the item, with the household level dropped. Appending j to index instead would ask one household to have bought item j both in kilograms and in bunches within one wave, which is precisely why the household key fails as an item key. The precedent for dropping the household is on the crop side: :func:_survey_median_factors groups (country, t, u, condition) and never carries i.

Two floors gate the item arm, and they gate different things: min_reports is the number of step-2 rows behind a (t, item, u) estimate (the twin of :data:SURVEY_MEDIAN_MIN_REPORTS), and min_baseline is the number of kg-known rows behind the (t, item) price-per-kg reference. The second is the one that binds. Neither applies to the item_col=None arm, which is ungated exactly as it was before #850 -- it is the FALLBACK rung, and gating a fallback would leave rows with nothing.

min_reports is applied PER WAVE, to each (t, item, u) estimate, and a wave that cannot clear it contributes nothing to the cross-wave median. It was applied to the support SUMMED over waves until the GH #850 red team measured it (2026-09-10), which meant one report per wave in five waves cleared a floor that four reports in a single wave did not. The crop side's :func:_survey_median_factors gates per wave; so does this now.

A THIRD gate exists and is OFF by default: baseline_max_spread refuses a (t, item) baseline whose own reports disagree by more than the given factor, measured as p90/p10 where the cell has 10 or more reports and max/min below that. The strict rung is otherwise ungated, and measured on the corpus it admits cells whose reports differ by up to 9e7x (median max/min 7.5 in Uganda, 178 in Malawi, 500 in Ethiopia). Left off pending a measured default; see the sweep in .coder/ledger/850-food-kg-item-axis-impl.md.

WHAT THE ITEM ARM IS AND IS NOT. It is exactly as good as the item's own kg-labelled rows. Where those are sound it is right (Niger's cowpea comes back at 1.11 kg per tiya against an answer key of 1.17); where they are not, it is wrong AND LOCALISED, whereas the u-pooled factor it replaces was wrong everywhere and diluted. Niger's millet is the worked example: nine Kg rows in one wave price millet at 2,286 FCFA/kg against a retail 250-450, and the item arm hands the resulting 0.33 kg/tiya to all 10,107 of that item's tiya rows. Removing a dilution is not the same thing as removing an error, and a movement is not a correction until something outside the estimate says so.

Parameters:

Name Type Description Default
df DataFrame

Food-acquired frame with Expenditure and Quantity columns and u (or unit_col) in the index.

required
price list[str]

Column(s) interpreted as expenditure for the ratio calculation.

['Expenditure']
quantity str

Column interpreted as quantity.

'Quantity'
index list[str]

Groupby levels for the per-item/period median step. At derive time this is ['t', 'i'] -- m is minted by _add_market_index inside _finalize_result, which runs AFTER the derived transform.

['t', 'm', 'i']
unit_col str

Name of the unit index level; renamed to u if different.

'u'
item_col str or None

Index level naming the item. None keeps the pre-#850 keys and the dict[str, float] return. 'j' keys the result on (item, unit).

None
min_reports int

Floor on the step-2 support behind a (t, item, u) estimate.

:data:`FOOD_KG_MIN_REPORTS`
min_baseline int

Floor on the kg-known rows behind a (t, item) baseline.

:data:`FOOD_KG_MIN_BASELINE`
min_baseline_tight

The dispersion-gated exception to min_baseline: a baseline with between min_baseline_tight and min_baseline - 1 reports is accepted when the max/min of its per-report price per kilogram is at or below tight_tolerance. See :data:FOOD_KG_TIGHT_TOLERANCE for why the spread is a range and not an IQR.

FOOD_KG_MIN_BASELINE_TIGHT
tight_tolerance

The dispersion-gated exception to min_baseline: a baseline with between min_baseline_tight and min_baseline - 1 reports is accepted when the max/min of its per-report price per kilogram is at or below tight_tolerance. See :data:FOOD_KG_TIGHT_TOLERANCE for why the spread is a range and not an IQR.

FOOD_KG_MIN_BASELINE_TIGHT
baseline_max_spread float or None

When given, a (t, item) baseline whose reports disagree by more than this factor is REFUSED however many of them there are -- p90/p10 at N >= 10, max/min below it. Off by default; the tight rung is unaffected (its own tolerance is stricter).

None

Returns:

Type Description
dict[str, float] or dict[tuple[str, str], float]

Mapping of unit label -> inferred kg factor when item_col is None; of (item, unit) -> factor when it is given. Keys are the RAW u label with its case PRESERVED -- the grouping is on the raw label, so Calebasse and calebasse come back as two entries with two factors. (The docstring used to say "(lowercased)", which was false; :func:_get_kg_factors is where the lower-casing happens, and it reports the resulting clashes -- see :class:UnitLabelCollisionWarning.) Units already in :data:KNOWN_METRIC, units in :data:_CURRENCY_DENOMINATED_UNITS, or units that cannot be inferred are absent from the output.

With _detail=True the item arm returns the underlying frame instead -- indexed (item, unit) with columns kg_per_unit, support, n_waves, wave_spread and baseline_tight. Private: it exists so :func:food_kg_factors can report WHICH rung, WHICH baseline and HOW MANY WAVES served a row without estimating twice.

crop_diversity

crop_diversity(crop_production, *, weight='count', value_col='Value_sold')

Shannon crop-diversity index per household (EPAR sdi).

METHODOLOGY transform. H = -Sigma_c p_c ln p_c over the crops a household grew, where p_c is crop c's share of whatever weight says the shares are shares OF.

Parameters:

Name Type Description Default
crop_production DataFrame

crop_production item feature, with a crop level named j or crop (resolved by :func:_resolve_crop_level).

required
weight (count, value, area)

What the shares are taken over.

  • 'count' -- the only basis available in every country. p_c is crop c's share of the household's crop OCCURRENCES, an occurrence being one distinct combination of the land and season levels the table carries: (plot, season) in Uganda, (plot,) in Malawi, falling back to raw rows only if the table has neither. De-duplicating to that grain is load-bearing, not tidiness: crop_production carries one row per (u, condition) as well, so counting raw rows would score a maize crop reported in two conditions as twice the maize.
  • 'value' -- p_c is crop c's share of the household's reported Value_sold. Covers SOLD output only, so a household that sold one crop and ate three scores H = 0: maximally specialised on the sales margin, which is a real fact about the household but is NOT its cropping diversity. Use it as a marketing-concentration measure, not as a substitute for the area basis.
  • 'area' -- raises :class:NotImplementedError. See below.
'count'
value_col str

Column used when weight='value'.

'Value_sold'

Returns:

Type Description
DataFrame

One float Crop_diversity column indexed by (t, i), non-negative, in NATS (natural log). A household growing one crop scores exactly 0.

Raises:

Type Description
NotImplementedError

For weight='area'. The area basis needs per-crop area PLANTED, and crop_production carries no such column in any country -- no AreaShare, no Area_planted (LEARNINGS.org L8 and the taxonomy's item-gap list record this). plot_features.Area is the whole plot's area, not crop c's share of it, and splitting it evenly across the plot's crops would MANUFACTURE the shares the index is made of -- on an intercropped plot the even split is the maximum-diversity answer by construction. The fix is a column (per-crop planted area or share), not a transform.

Notes

Sign. EPAR's sdi is Sigma p ln p -- it never negates (Uganda UNPS Wave 5/EPAR_UW_Uganda_UNPS_W5.do:3603,3617), so its published index is NON-POSITIVE and equals -H. Ours is the conventional Shannon H >= 0. Compare as sdi == -Crop_diversity, or a reader will think one of the two is broken.

Basis. EPAR weights by area PLANTED (area_plan), dropping rows with zero planted area -- which, its own comment notes, silently excludes permanent/tree crops unless they are a plot's only crop. That weight is refused here rather than approximated (see weight='area' above).

What one "occurrence" is, stated exactly, because it is easy to misread. The 'count' basis takes shares over distinct (t, i, crop, <plot level>, season) tuples -- so a crop grown on three plots DOES contribute three terms, exactly as in EPAR's plot-crop row basis. What is de-duplicated away is only the u / condition multiplicity WITHIN one land-season unit (a maize harvest reported in two conditions is one occurrence, not two). It is NOT an index over distinct crops: a household growing maize on four land-season units and beans on one scores H = 0.5004, not ln 2 = 0.6931. Pinned by tests/test_863_transforms.py:: test_crop_diversity_count_separates_plots_and_seasons.

Not comparable across countries on the 'count' basis. The occurrence key uses whichever of the land and season levels the country's crop_production carries -- (plot, season) for Uganda, (plot,) for Malawi, which has no season level at all. A Uganda household with an asymmetric crop mix across its two seasons therefore scores differently from an otherwise identical Malawi household, purely because one instrument records a season and the other does not. On a cross-country :class:~lsms_library.feature.Feature frame, compare within a country, or use weight='value' (whose shares have no land or season key) and accept its sold-only bias.

dependency_ratio

dependency_ratio(household_roster, *, working_age=(15, 64), age_col='Age')

Household dependency ratio (WB hh_dependency_ratio).

MECHANICAL reduction (below-the-line, GAP_RANKING Area 3). Per household: dependents / working-age members, where dependents are members younger than the working-age band's lower bound or older than its upper bound.

Parameters:

Name Type Description Default
household_roster DataFrame

household_roster item feature with an Age column, indexed by at least (t, i, pid).

required
working_age (int, int)

Inclusive [lo, hi] working-age band. Members with lo <= Age <= hi are the denominator; everyone else is a dependent. The default 15-64 matches the standard demographic definition the WB uses.

(15, 64)
age_col str

Age column name (case-insensitive lookup falls back to lowercase).

'Age'

Returns:

Type Description
DataFrame

One float Dependency_ratio column indexed by (t, i). Households with zero working-age members yield inf (all dependents, no earner) and are DROPPED — the WB column is likewise undefined there; an analyst wanting them kept reindexes afterward.

Notes

Reproduces hh_dependency_ratio (e.g. ETH_ESS1.do:957-961) up to the age-bucketing precision of Age. age_handler may return a fractional Age when DOB is available, so the band test uses < / > on continuous ages, not integer bins.

dummies

dummies(df, cols, suffix=False)

From a dataframe df, construct an array of indicator (dummy) variables, with a column for every unique row df[cols]. Note that the list cols can include names of levels of multiindices.

The optional argument =suffix=, if provided as a string, will append suffix to column names of dummy variables. If suffix=True, then the string '_d' will be appended.

farm_size

farm_size(plot_features, *, area_col='Area')

Total cultivated/owned plot area per household (WB farm_size).

MECHANICAL reduction (GAP 6 below-the-line). Σ plot Area per household over plot_features.

Parameters:

Name Type Description Default
plot_features DataFrame

plot_features item feature, grain (t, i, plot_id, ...), with a plot area column.

required
area_col str

Area column to sum.

'Area'

Returns:

Type Description
DataFrame

One Farm_size column indexed by (t, i) — total area in the plot_features.AreaUnit (Uganda: hectare-equivalent).

Notes

WB stores farm_size repeated per plot on the Plot dataset; this returns the HH total once. Reindex/broadcast to plot grain to compare directly. Plots with NaN/zero area drop from the sum.

fcs

fcs(food_acquired, *, groups, days, weights=None, cap=7)

WFP Food Consumption Score (EPAR fcs).

METHODOLOGY transform. Sigma_g weight[g] x min(days consumed in group g, 7) per household, over the eight WFP food groups.

Parameters:

Name Type Description Default
food_acquired DataFrame

food_acquired item feature (see :func:hdds).

required
groups Mapping[str, Sequence[str]]

Required. {group: [j label, ...]} naming exactly the keys of weights (default: the eight :data:FCS_WFP_WEIGHTS groups) and covering every j in the frame. Same partition rules as :func:hdds.

required
days str

Required, and there is deliberately no default. Name of the column giving the number of days in the past 7 that the item was consumed.

No country's food_acquired carries such a column today. The canonical schema (lsms_library/data_info.yml, Columns: food_acquired) declares Quantity / Expenditure / Price and nothing else, and the only Days column anywhere in the corpus belongs to food_coping. So fcs is UNUSABLE on a shipped table until a consumption-frequency column is wired, and that is the honest state of it -- an FCS is a frequency score, and deriving frequency from an acquisition amount would be inventing data. Where the source has the question (Tanzania's hh_sec_j3.hh_j08, which EPAR reads at W5.do:2574, is currently unwired on our side), wire it as a column and pass its name here.

required
weights Mapping[str, float]

Group -> weight. Defaults to :data:FCS_WFP_WEIGHTS.

None
cap int

Per-group day cap. A group's item-days are summed, then capped: eating rice on 5 days and maize on 4 is 7 staple-days, not 9.

7

Returns:

Type Description
DataFrame

One float FCS column indexed by (t, i).

Notes

Follows WFP's Technical Guidance Sheet (2008): sum the days within each group, cap the group total at 7, weight, sum. EPAR's Tanzania implementation diverges from that and the divergence is worth naming: rather than summing cereals and tubers and capping, it blends them (W5.do:2579-2582) as 7 if either is 7, else their sum if either is 0, else (max + min(sum, 7)) / 2 -- an averaging rule with no citation in the sheet. It also leaves its item-group 2 unweighted, so that group contributes nothing to the score at all. Neither behaviour is reproduced here.

The score is NOT thresholded into WFP's poor / borderline / acceptable bands (21 / 35) by this function -- the same separation as :func:rcsi / :func:rcsi_phase, so an analyst can apply the oil-and-sugar-adjusted cutoffs (28 / 42) where the country's diet warrants them without recomputing the score.

Unlike :func:rcsi, a group with no rows for a household scores 0 days rather than dropping the household: in a 7-day frequency recall, "no row" is the instrument's own zero. That is only true if the country's module really is a consumption recall over the whole item list -- see :func:hdds's divergence note 2.

fertilizer_rate

fertilizer_rate(plot_inputs, plot_features, *, nutrient='N', area_col='Area', volume_as_mass=True, on='parcel')

Fertilizer applied per unit plot area (EPAR "rate of fertilizer application").

MECHANICAL reduction. Divides a per-plot fertilizer mass by that plot's area, on the same land grain -- the same grain-bridge shape as :func:yield_kg.

Parameters:

Name Type Description Default
plot_inputs DataFrame

plot_inputs item feature (see :func:nitrogen_kg): an input level, a Quantity and a unit u.

required
plot_features DataFrame

plot_features item feature, grain (t, i, plot_id, ...), carrying Area -- hectares, per the canonical schema ("plot area in hectares (canonical unit)", lsms_library/data_info.yml, Columns: plot_features: Area). So the returned rate is kg/ha wherever the country honours that.

required
nutrient (N, product)
  • 'N': numerator is :func:nitrogen_kg -- kilograms of NITROGEN, i.e. product kg times the nutrient share of :data:_NITROGEN_CONTENT. Output column Nitrogen_kg_per_ha.
  • 'product': numerator is raw fertilizer PRODUCT kg -- the same rows, same unit conversion, N-share forced to 1.0. Output column Fertilizer_kg_per_ha. Note this covers the INORGANIC vocabulary only: _NITROGEN_CONTENT's keys are the nutrient classes (nitrate / phosphate / potash / mixed) and the product names (urea, CAN, SA, DAP, NPK, TSP, SSP, MOP) -- manure and other organic inputs carry no key and contribute nothing to either basis.
'N'
area_col str

Plot-area column in plot_features.

'Area'
volume_as_mass bool

Forwarded to the unit->kg conversion.

True
on (parcel, plot)

Land grain to join the two features on. The default matches :func:yield_kg's, and that agreement is deliberate: the two transforms bridge the same two plot vocabularies, and a pair of sibling functions that default differently is a trap. It was one -- with on='plot' as the default, fertilizer_rate(Uganda) returned an EMPTY frame and said nothing (see :class:PlotGrainMismatchWarning).

  • 'parcel' (default): reconcile the two plot vocabularies to their common parcel key first, exactly as :func:yield_kg does (_parcel_from_crop_plot / _parcel_from_feature_plot, both no-ops on a vocabulary that carries no {hhid}- prefix or _suffix). Uganda needs it: {hhid}-{parcel}-{plot} on the ag modules vs. {parcel}_{suffix} on plot_features, which share no literal key at all -- 0 rows on 'plot', 509 on 'parcel'. Measured no-op on Malawi, whose two features already share the literal key (66,298 of 66,368 distinct (t, i, plot_id) keys match, 99.9%, identically under either setting).
  • 'plot': join the literal plot key verbatim. Use where the two features are known to share a vocabulary and the parcel extraction would be wrong.
'parcel'

Returns:

Type Description
DataFrame

One column -- Nitrogen_kg_per_ha or Fertilizer_kg_per_ha, named for the numerator so the basis cannot be lost -- indexed by (t, i, plot_id) (or (t, i, parcel) for on='parcel').

Notes

Divergences from EPAR's rate, both in the denominator and the numerator, and neither is a bug on either side:

  • Denominator. EPAR's is hectares planted, itself imputed as parcel area times a reported planted share ("Plot area size = Parcel area size x (planted share)", its readme). Ours is the plot area the instrument reported and the country converted to hectares (plot_features.Area, required). A plot only partly planted therefore gets a LOWER rate here than in EPAR's table.
  • Numerator. EPAR's rate is over fertilizer PRODUCT; the WB harmonised panel's nitrogen_kg is over the nutrient. Hence nutrient=, with no default that hides which one you got.
  • Inherits :func:nitrogen_kg's unit-conversion coverage caveat (only metric-named u rows convert) and its class-nominal N shares, so on a country recording nutrient class rather than product the rate is a nominal, not a measured, nitrogen figure.

Plots with fertilizer but zero or missing area are dropped (an infinite rate is not a rate).

food_acquired_valued

food_acquired_valued(df, valuation, *, geo=None, threshold=10, value_col='Expenditure', quantity_col='Quantity')

Value food_acquired's unvalued non-purchased rows, with PROVENANCE.

METHODOLOGY transform (GH #585). The item-grain half of food_expenditures(basis='total', valuation=...), exposed on its own so the imputation can be AUDITED row by row -- which rung served each row, and at what unit price. Every row of df gets a row here, on the same index and in the same order.

This is opt-in and it is never the default, for a measured reason. Every rung below prices own-consumption at a PURCHASE transaction, and a purchase sits above the farm gate by the marketing margin. Our five-source price study measured the gap: in Uganda, own-consumption valued at the purchase price sits about 23% above what a sale actually fetched (slurm_logs/price_sources/SYNTHESIS.org:750-753; the transitive 223-cell set, ratio 1.234). That number is Uganda-only -- GhanaLSS ships no sale-price source, so no comparable figure exists for it. The UNPS interviewer manual is explicit that column 9 "should be valued at farm gate/producer price ... excludes any cost transport and marketing services" (SYNTHESIS.org:732-736), so a consumption aggregate built this way carries part of a margin the instrument intended to exclude. Deaton & Zaidi (2002, Guidelines for Constructing Consumption Aggregates, LSMS WP 135) treat the producer/market choice as the two defensible readings and warn that the market one inflates the aggregate relative to purchases. Choose it deliberately; do not reach for it because basis='total' looked incomplete.

Rungs, in precedence order (ValuationSource records which one served)

reported The row already carries a non-zero value_col. Never re-valued -- including a value the survey itself imputed for s='inkind' (lsms_library/data_info.yml:602). Zero counts as missing, exactly as :func:food_expenditures_from_acquired's own replace(0, nan) does, so a zero-valued produced row IS a candidate. own_price The same household's own purchase unit value for the same (t, j, u): the sum of that household's purchased value_col over the sum of its purchased quantity_col, in that wave, for that item, in that unit. No cross-household information whatever. This is EPAR's consumption repo's FIRST rung (EthiopiaW5_...do:423, UgandaW4_...do:693), conditional there too on the consumed unit matching the purchased one. Note that EPAR's Ag repo does the opposite -- it keeps the household's own price as a parallel value_harvest_hh series rather than a rung -- so "EPAR's ladder" is ambiguous and this docstring says which (LEARNINGS.org L5 item 4). median_price The geographic median-price ladder, run by :func:median_price_valuation on a price pool built only from s == 'purchased' rows (value_col / quantity_col). The ladder walks geo finest -> coarsest and takes the median of the finest cell with at least threshold priced observations, with an unconditional national fallback. Cells are (geo, t, j, u): t is in the item key because median_price_valuation has no wave axis of its own, and without it the national rung would pool currency across a decade. none No rung produced a price. Counted, never hidden. Also covers a purchased row with no recorded outlay, which is a data defect rather than an unvalued acquisition and is deliberately not a candidate.

Parameters:

Name Type Description Default
df DataFrame

food_acquired, canonical index (t, v, i, j, u, s) (data_info.yml:51); v optional. Needs s as an index level or column -- without an acquisition source there is nothing to value.

required
valuation str or sequence of str

Rungs to run, IN THE ORDER GIVEN; each fills only rows still unvalued. A scalar means exactly that rung -- 'median_price' runs the spatial ladder alone, it does NOT quietly run 'own_price' first. The composed form ('own_price', 'median_price') is EPAR's consumption-repo ordering and is what you want if you want theirs.

required
geo DataFrame

Coarser geographic rungs, indexed by (t, v) -- i.e. cluster_features' own grain -- with columns ordered finest -> coarsest (e.g. [['District', 'Region']]). The v level of df's own index, when present, is prepended as the finest rung. With geo=None the ladder is v (if present) then national.

None
threshold int

Minimum priced observations for a geographic cell's median to be adopted; forwarded to :func:median_price_valuation (the WB / EPAR-Ag gate of >= 10).

10
value_col str

Column names on df.

'Expenditure'
quantity_col str

Column names on df.

'Expenditure'

Returns:

Type Description
DataFrame

Indexed like df, with

Expenditure (i.e. value_col) the reported value where there was one, the imputed value otherwise, NaN where no rung could serve. ValuationSource one of :data:VALUATION_SOURCES. valuation_price the unit price the winning rung used, per the row's own u (NaN for reported and none). Deliberately NOT called Price: that name is the survey's own reported unit price, and this one is CONSTRUCTED.

Two records ride on .attrs:

valuation_sources {source: n_rows} over the INPUT rows. The four :data:VALUATION_SOURCES partition the frame and sum to len(df). The extra candidates key counts the rows that were ELIGIBLE for valuation (non-purchased, no reported value) and is deliberately outside that partition, since such a row is still served by one of the four. POOLED on a cross-country :class:~lsms_library.feature.Feature frame -- read it as a total, never as coverage. valuation the tuple of rungs actually run, and valuation_geo_levels the ladder they ran on.

Notes

The price basis is the row's NATIVE unit, not kilograms. u is an item key, so a produced row is valued at the purchase price of the very label it was reported in and unit alignment is exact by construction. A kg-normalised basis would pool more finely-split labels (Malawi carries Kilogramme, Kilogram, Kg and kg as four labels in one wave) but would inherit the open kg-inference defect of GH #850, and it buys nothing measurable: national-rung coverage at threshold=10 is 94.2% (Malawi) / 76.5% (Uganda) on the native key, and lower-casing u moves it 0.00pp in both. A u='Value' row prices at 1 currency-unit per currency-unit, which is the right answer for it.

Read that 94.2% as COVERAGE, not as accuracy, and read the value-weighted complement beside it. median_price_valuation's national rung is an unconditional fallback, so threshold gates the geographic cells and gates nothing at the top of the ladder. Measured on Malawi (red-team, 2026-09-09), by share of the imputed MONEY rather than of the rows: 6.48% of the imputed value is priced off fewer than ten purchase observations nationwide, and 2.85% off exactly one. The 5.8% of candidate rows below the gate are not a random 5.8% of the money. The concrete case: a single 2016-17 purchase sets "Small Animal - Rabbit, Mice, Etc." (u='Whole') at 100,000 MWK, which is then applied to 195 rows and contributes 2.12% of Malawi's entire delta. The median pool is 533 observations, so the bulk is well supported -- but a thin (t, j, u) cell is thin for every household in it, so this tail is systematic rather than an outlier. It is COUNTED here, never clipped; valuation_price is returned per row so a caller can gate on it.

Whether the national rung should carry a gate of its own is a follow-up, deliberately not decided here. @ligon has ruled on the parallel food-side inference (GH #850) that a thin baseline is a floor of 5 with a dispersion-gated exception at 3-4; adopting the same rule at this rung is the obvious candidate, and it would change returned numbers, so it belongs in its own change.

Nothing here is cached. food_expenditures is derived at read time (Country._FOOD_DERIVED) and this runs inside that derivation, so no parquet ever holds an imputed value.

food_expenditures_from_acquired

food_expenditures_from_acquired(df, basis='purchased', *, valuation=None, geo=None, threshold=10)

Derive food expenditures from food_acquired.

Returns a DataFrame of expenditure per household × item × period × acquisition source (when s is present in the input index), summed over units.

Parameters:

Name Type Description Default
df DataFrame

food_acquired with an Expenditure column.

required
basis (purchased, total)

Which acquisition sources contribute to Expenditure (GH #575):

  • 'purchased' (default): cash outlay only — keep rows with s == 'purchased'. Own-production / in-kind / other sources carry no cash expenditure, so they are excluded. This is consistent across countries regardless of whether the source happened to record an imputed produced/in-kind value: some waves (e.g. Serbia, GhanaSPS) populate Expenditure on produced/ in-kind rows, others (e.g. Guatemala, and every country built via the stock food_acquired_to_canonical) leave it NaN. Filtering to purchased removes that cross-country divergence.
  • 'total': all recorded acquisition value — sum Expenditure across every s. Where the source recorded a produced/in-kind value it is included; where it did not, 'total' equals 'purchased' for that country -- unless valuation= is passed, which is the ONLY way this function ever fabricates a value.
'purchased'
valuation str or sequence of str

Opt-in own-production / in-kind valuation (GH #585). None (the default) is today's behaviour exactly: no value is fabricated. Otherwise the rungs named are run, in the order given, over the non-purchased rows that carry no value -- see :func:food_acquired_valued for the rungs, the provenance, and the measured purchase-side bias, which you should read before using this: every rung prices own-consumption at a purchase transaction, and in Uganda own-consumption so valued sits about 23% above what a sale actually fetched (slurm_logs/price_sources/SYNTHESIS.org:750-753, Uganda-only), while the interviewer manual asks for the farm-gate price (:732-736). Requires basis='total'; under basis='purchased' there is no non-purchased row left in the output to value, so it raises rather than silently doing nothing.

A scalar names EXACTLY that rung: 'median_price' runs the spatial ladder alone and does not quietly run 'own_price' first. Pass ('own_price', 'median_price') for EPAR's consumption-repo ordering.

None
geo DataFrame

Coarser geographic rungs for the 'median_price' ladder, indexed by (t, v) with columns ordered finest -> coarsest. See :func:food_acquired_valued.

None
threshold int

Minimum priced observations for a geographic cell's median to be adopted. Forwarded to :func:median_price_valuation.

10
Notes

Phase 4 of GH #169 preserves the s (acquisition-source) level in the output. Users who want the legacy collapsed view call food_expenditures.groupby(level=['t','v','i','j']).sum() explicitly; basis='purchased' then yields the purchased total, 'total' the sum across recorded sources.

When the input has no s level (pre-canonical waves) the two bases coincide — there is no source split to filter on.

With valuation=, the returned frame carries the per-source tallies on .attrs['valuation_sources'] (plus 'valuation' and 'valuation_geo_levels'), mirroring harvest_kg's attrs['kg_factor_sources']. The per-row ValuationSource column is NOT on this frame: the output is summed over u, and one (t, i, j, s) cell can mix rungs, so a per-row label has nowhere unique to land. Call :func:food_acquired_valued directly for the item-grain frame that carries it.

A :class:~lsms_library.feature.Feature call forwards valuation= (it forwards by signature) and its VALUES are exact -- Malawi sums to the same 426,562,493.98 either way. Its attrs depend on how many countries were asked for, and the boundary is worth stating precisely because a reader who tests the general claim on one country would see it contradicted:

  • one country -- a concat of a single frame, nothing to disagree with, so the tallies COME THROUGH;
  • more than one -- the frames' attrs disagree by construction, one record per country, which lands in the {} row of the propagation rule (CLAUDE.md, "Panel ID Transitive Chains"), so attrs['valuation_sources'] is ABSENT.

Pooling the tallies across countries is deliberately NOT built: a pooled median_price: 280623 would say nothing about WHICH country was imputed, the same objection harvest_kg's docstring makes about its own pooled counts. Ask per country, or group the item-grain frame.

food_kg_factors

food_kg_factors(df, *, volume_as_mass=True, item_col='j', unit_kg=None, min_reports=FOOD_KG_MIN_REPORTS, min_baseline=FOOD_KG_MIN_BASELINE, min_baseline_tight=FOOD_KG_MIN_BASELINE_TIGHT, tight_tolerance=FOOD_KG_TIGHT_TOLERANCE, baseline_max_spread=FOOD_KG_BASELINE_MAX_SPREAD)

Per-ROW kg-per-unit for a food_acquired frame, with provenance.

The food twin of :func:harvest_kg_factors, and deliberately the same shape: one row per input row, a resolved kg_per_unit, a KgFactorSource drawn from :data:FOOD_KG_FACTOR_LAYERS, the per-layer candidates beside it, and counts in attrs['kg_factor_sources'] that PARTITION the frame.

Why per-row at all (GH #850): once the inference carries an item axis, a factor is no longer a property of the unit label, so there is no dict to return. And the ladder is now worth reporting -- a row served by its own (j, u) cell and a row served by the u-pooled fallback were indistinguishable before, and after this change they differ by more than they did before it.

Parameters:

Name Type Description Default
df DataFrame

food_acquired-shaped, with u in the index and Quantity / Expenditure columns. Without Expenditure there is nothing to take a price ratio of, and only the survey_kg / metric layers can serve a row.

required
item_col str

The item level. Absent from the frame -> the item rungs serve nothing and every inferred row falls to unit; the layers still partition.

``'j'``
volume_as_mass

Passed through; see :func:conversion_to_kgs.

True
min_reports

Passed through; see :func:conversion_to_kgs.

True
min_baseline

Passed through; see :func:conversion_to_kgs.

True
min_baseline_tight

Passed through; see :func:conversion_to_kgs.

True
tight_tolerance

Passed through; see :func:conversion_to_kgs.

True

Returns:

Type Description
DataFrame

Indexed like df, with columns kg_per_unit, KgFactorSource, kg_survey, kg_metric, kg_item_unit, kg_item_n_waves, kg_item_wave_spread, baseline_refused_spread and kg_unit. attrs['kg_factor_sources'] maps each layer to its row count; attrs['kg_factor_wave_spread'] reports how far the pooled waves disagreed; attrs['kg_factor_baseline_gate'] reports how many rows the dispersion gate cost their own (item, unit) factor and where they landed instead (rows_refused == refused_to_unit + refused_to_none). Only the first partitions the frame.

Notes

WHAT THIS DOES NOT FIX, stated where a user will meet it. The only external answer key the corpus has for a food unit is Mali's own questionnaire, and against it the corrected inference is BETTER and still WRONG: for a stated 100 / 50 / 25 kg sack it serves a median 50.0 / 34.4 / 14.1 kg per row (0.50x / 0.69x / 0.56x of truth, up from 0.33x / 0.52x / 0.21x). Gramme is now exact, but only because GH #850 taught the label parser to READ it -- no inference was involved. Mali's CONTENTS.org advice ("use native units, or u = 'Kg'") therefore survives this change. A price-ratio inference cannot be argued into being a conversion table; where a country ships one, use it (the crop side's shipped layer, GH #852/#854).

kg_per_unit on a survey_kg row is the factor IMPLIED by the survey's own kilograms, Quantity_kg / Quantity, so that Quantity * kg_per_unit reproduces the delivered kilograms on every row rather than on all but those. Where that quotient cannot be formed (a missing or zero Quantity) the row keeps the ladder's factor and still reports survey_kg: the kilograms DO come from the survey, and the factor column is there for the price-per-kg conversion, which would otherwise lose a row it can serve.

food_prices_from_acquired

food_prices_from_acquired(df, units='kgvalue', *, volume_as_mass=True, unit_kg=None)

Derive food prices from food_acquired.

Returned at the canonical (t, i, j, u, s) grain (v is omitted and re-joined from sample() by _finalize_result at API time). Any extra recall level present in the input — notably GhanaLSS's visit — is aggregated out here via the median Price, mirroring how food_quantities_from_acquired / food_expenditures_from_acquired collapse it with .sum() (price is per-unit, not additive, so median rather than sum). This keeps the three _FOOD_DERIVED transforms at the same country-level grain so Feature('food_prices') no longer has to drop a leaked visit level (and silently keep one arbitrary visit's price) before the cross-country concat — see GH #517.

Parameters:

Name Type Description Default
df DataFrame

food_acquired DataFrame. Must have Quantity; needs Expenditure for the *value modes and Price for the *price modes.

required
units (kgvalue, kgprice, unitvalue, unitprice)

Which Price to compute, varying across two axes (denominator and source):

  • 'kgvalue' (default): Expenditure / Quantity_kg (currency / kg, derived). Backward-compatible with the pre-Phase-4 implementation.
  • 'unitvalue': Expenditure / Quantity (currency / native u, derived). For u='Value' rows (LCU-only goods) the formula gives 1 — Kwacha per Kwacha — mathematically correct but analytically useless; consumers should filter on u before aggregating.
  • 'kgprice': reported Price × kg_factor (currency / kg). NaN where Price is missing or u is unconvertible to kg.
  • 'unitprice': reported Price (currency / native u). NaN where the survey did not record a unit price. This mode reflects the canonical food_acquired.Price column — market price for s='purchased', farmgate for s='produced', imputed for s='inkind'.

See slurm_logs/DESIGN_food_prices_units_kwarg_2026-05-06.org.

'kgvalue'

Returns:

Type Description
DataFrame

Single-column Price DataFrame at the canonical (t, i, j, u, s) grain (v re-joined downstream; any extra recall level such as visit is collapsed via median Price). Zero / infinite / NaN prices are dropped.

Notes

The 'kgvalue' default deliberately departs from the term-of-art "unit value" common in the literature (e.g. Deaton 1988, 1997), which usually means Expenditure / Quantity standardized to kg. The 'kgvalue' / 'unitvalue' naming makes the denominator explicit at the cost of mild inconsistency with prior usage; consult the docstring before substituting one for "unit value" in literature review or reproduction work.

No silent fallback between modes. 'unitprice' returns NaN where Price is missing rather than falling back to 'unitvalue'; a caller wanting "best available" combines results explicitly with their own provenance tracking.

For canonical s-axis input, only s='purchased' rows have a meaningful Expenditure (the Phase-3 helpers set produced/inkind Expenditure to NaN). Under 'kgvalue' / 'unitvalue' those rows become NaN-after-divide and drop out. Under 'kgprice' / 'unitprice' produced/inkind rows survive iff the wave script populated the survey-reported Price column upstream (Uganda Phase 3 path). This function does not synthesize a Price.

food_quantities_from_acquired

food_quantities_from_acquired(df, units='kgs', *, volume_as_mass=True, unit_kg=None)

Derive food quantities from food_acquired.

Parameters:

Name Type Description Default
df DataFrame

food_acquired DataFrame with a Quantity column and a u index level naming the unit each row's quantity is in.

required
units (kgs, units)

Aggregation basis:

  • 'kgs' (default): convert Quantity to kilograms where the unit's kg factor is known (via :func:_get_kg_factors); tag those rows with u='kg'. Rows whose unit lacks a factor (e.g. u='Value' for LCU-only goods such as "meals in restaurants", or u='tin' when no per-tin kg conversion is known) are carried through with their native Quantity and original u label, NOT dropped. The output is therefore mixed-physical-unit; the u index distinguishes kg from native rows. Consumers wanting purely-kg rows do df.xs('kg', level='u').
  • 'units': sum native Quantity per (t, v, i, j, u, s), no kg conversion attempted.
'kgs'

Returns:

Type Description
DataFrame

Single-column Quantity DataFrame with u and s retained in the index (Phase 4 of GH #169 preserves the acquisition-source axis).

Notes

The carry rule for unconvertible units in 'kgs' mode is the Phase-4 design call recorded in slurm_logs/DESIGN_food_prices_units_kwarg_2026-05-06.org. It differs from the original implementation, which silently dropped unconvertible rows from food_quantities.

The output's group-by preserves both u (the per-row unit; 'kg' for converted rows, native otherwise) and s (acquisition source), per the GH #169 canonical schema. Pre-canonical waves where s is absent silently skip the s level.

format_interval

format_interval(interval, compact=True)

Render a pd.Interval as a human-readable column-name label.

Two styles, selected by the compact flag:

  • Compact (default) — matches the historical household_characteristics column names. Finite integer bounds span-1 apart or wider collapse to "{lo:02d}-{hi-1:02d}" (e.g. [0, 4) → "00-03"); the unbounded-top bucket renders as "{lo:02d}+" (e.g. [51, inf) → "51+").
  • Explicit — half-open interval notation "[lo, hi)" with "{lo}+" for the unbounded top bucket. Used whenever any bound is fractional (spans shorter than a year).

:func:roster_to_characteristics picks compact=True when all age_cuts are integers and False otherwise, so column names stay consistent within a single call.

gross_crop_revenue

gross_crop_revenue(crop_production, *, by=None, value_col='Value_sold')

Reported crop SALE revenue per household (EPAR value_crop_sales).

MECHANICAL reduction. Sums crop_production.Value_sold -- the sale value the household actually REPORTED -- to the (t, i) grain, or to (t, i, crop) with by='j'.

Parameters:

Name Type Description Default
crop_production DataFrame

crop_production item feature carrying a reported sale-value column. Grain varies by country: (t, i, plot_id, j, u, condition, season) in Uganda, (t, i, plot_id, crop, u) in Malawi.

required
by (None, j, crop)

None (default) reduces all the way to (t, i). Any of 'j' / 'crop' keeps the crop level, whichever name this country uses for it (resolved via :func:_resolve_crop_level, so by='j' works on Malawi's crop level too) -- the output level keeps the table's own name.

None
value_col str

Reported sale-value column to sum.

'Value_sold'

Returns:

Type Description
DataFrame

One Gross_crop_revenue column, in nominal local currency, indexed by (t, i) (or (t, i, <crop level>)).

Notes

This is NOT EPAR's "gross value of crop production", and the two must not be compared as if they were. EPAR's construct (Uganda UNPS Wave 5/EPAR_UW_Uganda_UNPS_W5.do, GROSS CROP REVENUE section) values the entire harvest -- including the share never sold, which for a smallholder panel is most of it -- at a median unit price imputed from observed sales on a geographic ladder. That is a METHODOLOGY transform, and the library already has it: :func:median_price_valuation, whose ladder (finest geographic cell with >= 10 priced observations, then a national fallback) is the same one EPAR and the WB both use. Compose the two -- value the harvest with median_price_valuation, then sum -- to get their quantity. This function deliberately reports only what the instrument wrote down.

No de-duplication, by design (count, never clip). Where a country attached a household-crop sale total to more than one physical row, a plain sum would double-count it. It does not: measured on the warm corpus (2026-09-09), at Uganda's finest key (t, i, plot_id, j, season) only 963 of 44,598 groups hold more than one non-zero sale row, and in just 37 of those do all rows share ONE Value_sold -- the stamping shape. 926 carry genuinely differing values, which is what the instrument implies: UNPS question 6 is compound (how much, in what unit, in what condition) and the sale questions ride the SAME record, so each condition row carries its own quantity-sold and value-sold (Uganda/_/CONTENTS.org §"Harvest condition", and the slot-2 column list s5bq07a_2 / s5bq08_2). Summing across u / condition / season is therefore correct, and 895 exact (t, i, crop) repeats out of 45,592 rows (1.9% of value) are ordinary coincidence -- two plots of one crop sold in equal quantity for equal money.

On Malawi the risk runs the OTHER way: this is a systematic LOWER BOUND, not a double count. Its sale is a HOUSEHOLD-crop question that data_scheme.yml records as "attached to single-plot crops only", and the consequence is measurable: of 87,440 (t, i, crop) groups grown on ONE plot, 23,206 (26.5%) carry a sale row, against 514 of 19,311 (2.7%) for crops grown on more than one plot. A crop spread over two or more plots is almost never credited with a sale at all, by instrument design -- so Malawi's Gross_crop_revenue UNDER-states sales, and the shortfall is concentrated in exactly the households farming the most land. Do not read a Malawi total as complete.

A household present in crop_production that reported no sale at all is ABSENT from the output, never 0 -- sum(min_count=1), the same never-zero-fill rule as :func:rcsi. Reindex against sample() if you want the non-sellers as explicit zeros.

harvest_kg

harvest_kg(crop_production, *, volume_as_mass=True, carry_native=False, min_reports=SURVEY_MEDIAN_MIN_REPORTS, shipped_factors=None)

Total harvested kilograms per (t, i, plot_id, j) from crop_production.

MECHANICAL reduction (GAP 1 → WB Plotcrop/Plot harvest_kg). For each reported harvest row, convert the native-unit Quantity to kilograms, then sum within each plot-crop.

Each row's kg-per-unit factor is built by a LAYERED procedure, in precedence order: the row's own reported KgFactor (what the instrument wrote down, e.g. UNPS a5?q6d); a shipped factor from an externally supplied conversion table, when the caller passes one as shipped_factors; the median of reported factors for the same (u, condition) in the same country-wave, where at least min_reports rows report one; the library's inferred factor from the shared unit→kg machinery (:func:_get_kg_factors via :func:_kg_factor_series — the same factors :func:food_quantities_from_acquired applies to food_acquired); and otherwise none, in which case the row contributes nothing. The survey's own number wins wherever it exists, on the same principle by which food_acquired's exact per-row Quantity_kg already beats an inferred factor.

Which layer served each row is DISCLOSED, not assumed: the counts ride on result.attrs['kg_factor_sources'] and the per-row provenance is available from :func:harvest_kg_factors, which also reports how often the survey's factor and the library's inferred one DISAGREE. A frame with no KgFactor column can only reach the inferred layer, so it gets exactly the result it got before the column existed.

Parameters:

Name Type Description Default
crop_production DataFrame

The crop_production item feature, indexed by (t, i, plot_id, j, u, season) (v may also be present — it is ignored for the reduction and dropped from the output grain) with a reported Quantity column in native unit u.

required
volume_as_mass bool

Forwarded to :func:_get_kg_factors; treat 1 litre = 1 kg for fluid units (juice/beer harvest rows) when True.

True
min_reports int

How many rows of a (u, condition) group must carry a reported KgFactor before their median is used to fill that group's unreported rows. Inert on a frame with no KgFactor column.

:data:`SURVEY_MEDIAN_MIN_REPORTS`
shipped_factors DataFrame

An externally shipped kg-per-unit table, forwarded verbatim to :func:harvest_kg_factors -- see there for its shape, the join keys, the ambiguity refusal, and why nothing auto-discovers it. Omitted (the default), this whole layer is inert: every Harvest_kg value is bit for bit what it was before the layer existed. (The counts dict gains three zero-valued keys and :func:harvest_kg_factors gains three columns -- forced by the partition invariant, and no number moves.) Example::

harvest_kg(
    cp,
    shipped_factors=ethiopia.crop_conversion_factors(...))
None
carry_native bool

When False (default, matching the WB construct), rows whose u has no known kg factor contribute NOTHING to the sum — Harvest_kg reflects only the convertible quantity, and a plot-crop with no convertible rows is absent from the result. When True, those rows are carried at their native Quantity (a deliberate over-count used only for sensitivity analysis); the default False is the parity-faithful choice.

False

Returns:

Type Description
DataFrame

One Harvest_kg column indexed by (t, i, plot_id, j) (whichever of those levels are present), summed over native units and season.

result.attrs['kg_factor_sources'] is {layer: n_rows} over the INPUT rows and sums to len(crop_production). It counts rows by the layer that gave them a FACTOR, not by whether they contributed: a reported row whose Quantity is 0 or NaN is counted reported and still drops out of the sum below. result.attrs['kg_factor_disagreement'] carries the layer-pair audit described in :func:harvest_kg_factors.

Notes

Divergence from WB: the WB Uganda code applies survey-provided conversion tables that cover bespoke local containers ("Sack (100 kgs)", "Basket (Unspecified)", regional "Heap"/"Bunch" sizes). Our shared factor map only resolves units that name their metric content (e.g. "Basket (10 kg)", "Nice cup (60g)") plus the hand-coded metric tokens, so on Uganda it converts ~21% of crop rows through that layer alone (28 147 of 133 683, measured 2026-09-09). Harvest_kg is therefore a lower bound on the WB figure for plots dominated by non-metric containers; magnitudes agree on metric-reported plots.

Closing that gap is a per-country job in three forms, none of them a change to this transform: extend the country's u-table (harvest_units); wire the instrument's own conversion factor into the optional KgFactor column, which the reported and survey_median layers above then consume; or -- where the survey ships a conversion TABLE rather than a per-row factor -- load that table and pass it as shipped_factors. The third is the one Ethiopia and Malawi need (GH #852, #854): their tables are keyed on crop x unit x region, so they can never be a per-row column, and the de-duplication and region resolution they require are decisions that belong at the call site rather than baked into a parquet. The second reaches rows the first cannot -- Uganda's 2018-19 season-A harvest side ships a 100%-populated a5aq6d conversion factor and NO harvest-unit code at all, so its 7 153 rows sit at u='Unknown' and no unit table can ever convert them.

The survey_median layer REFUSES the missing-unit sentinel, on both sides. u='Unknown' (:data:U_UNKNOWN) is the absence of a unit, not a unit, so a "group" of such rows pools containers of unrelated sizes: a median over them is a number about nothing, and filling other such rows with it would fabricate weights. A row with no unit recorded therefore gets kilograms from its OWN reported KgFactor or from nothing at all -- which is why wiring KgFactor is the only thing that can ever convert Uganda's 2018-19 season-A rows.

harvest_kg_factors

harvest_kg_factors(crop_production, *, volume_as_mass=True, min_reports=SURVEY_MEDIAN_MIN_REPORTS, shipped_factors=None)

Per-row kg-per-unit factor for crop_production, and its PROVENANCE.

The factor half of :func:harvest_kg, exposed on its own so the layers can be AUDITED -- most usefully, the survey's own reported weights against the library's inferred ones for the same unit. Every row of crop_production gets a row here, in the same order and on the same index.

Layers, in precedence order (KgFactorSource records which one served each row)

reported The row's own KgFactor column -- what the instrument wrote down (UNPS a5?q6d, "conversion factor into kg"), where it is finite, > 0 and PLAUSIBLE (:func:_screen_reported_factors: a kilogram unit must weigh 1, nothing may exceed :data:KG_FACTOR_MAX, and an integral calendar year on a non-kg unit is the adjacent-column signature). A rejected value is never clipped or rescaled -- the row falls through to the layers below exactly as if it had not reported, and the rejection is counted under reported_implausible. KgFactor is REPORTED, never constructed: see the canonical schema note in lsms_library/data_info.yml. shipped A factor looked up in shipped_factors -- an EXTERNALLY SHIPPED conversion table the caller passed in (the World Bank's Ethiopia Crop_CF_Wave{2..5}, Malawi's IHS5 crop files: GH #852, #854). Screened by the SAME :func:_screen_reported_factors, with its rejections counted under shipped_implausible; a rejected shipped factor is likewise never clipped, and the row falls through. Ranked below reported because a table is not this row's enumerator, and above survey_median because it is external evidence rather than a pooling of the survey's own reports. See shipped_factors below for the table's shape and the join. survey_median The median reported KgFactor of the same (u, condition) within the same country-wave, where at least min_reports rows report one; falling back to u alone. The survey's own weights filling the survey's own gaps -- see :func:_survey_median_factors. inferred The shared unit→kg machinery, unchanged: :func:_kg_factor_series → :func:_get_kg_factors (hand-coded metric tokens, the explicit-metric label parser, then price-ratio inference). This is the ONLY layer that acts on a frame with no KgFactor column, so such a frame gets exactly the factors it got before this function existed. none No layer produced a usable factor. Counted, not hidden: these rows contribute nothing to Harvest_kg.

Parameters:

Name Type Description Default
crop_production DataFrame

The crop_production item feature. Needs Quantity only for the price-ratio branch of the inferred layer; needs u as an index level or a column. KgFactor, condition and country are all optional.

required
volume_as_mass bool

Forwarded to :func:_get_kg_factors (1 litre = 1 kg for fluids).

True
min_reports int

N for the survey_median layer. One number is applied to every group, but the groups are per country-wave, so a thin module does not borrow a thick one's licence.

:data:`SURVEY_MEDIAN_MIN_REPORTS`
shipped_factors DataFrame

An externally shipped kg-per-unit table. KgFactor column (kg per ONE unit of the row's u) plus an optional Source string, keyed -- as index levels or as columns -- on some subset of :data:SHIPPED_FACTOR_JOIN_LEVELS ('country', 't', 'j', 'crop', 'crop_variety', 'u', 'condition', 'region'). The loader decides which levels its table can honestly key on; the join uses those present in BOTH the table and crop_production. j and crop are both there because the corpus spells the crop level both ways (Uganda/Ethiopia/Tanzania j; Malawi and nine others crop); key the table on whichever name its own country's frame carries. crop_variety is the un-collapsed crop, for a table keyed finer than the served label.

NOTHING AUTO-DISCOVERS A COUNTRY'S TABLE, and the result is never stored. The intended call shape, once the loaders of GH #852 / #854 exist, is explicit::

harvest_kg(cp, shipped_factors=ethiopia.crop_conversion_factors(...))

A shipped_factors='auto' registry is a live design question for @ligon; the options and the trade-off are written up in .coder/ledger/harvest-kg-shipped-factors.md, section 6.

region must be resolved by the CALLER onto crop_production (join cluster_features and rename its Region to region); this function never re-enters sample() or cluster_features(). Note also that crop_production's crop level (j or crop) is the DECODED crop label, not the WB crop code -- a table read straight from Crop_CF_Wave4.dta carries 74, which matches no j, and the shipped_matched: 0 count is the only tell. Decode in the loader. Malawi's IHS5 tables need one further step: their crop key is a VARIETY ('MAIZE HYBRID', 'RICE FAYA') and the served crop is the aggregate Preferred Label, so the loader must collapse the varieties -- and refuse the keys whose varieties disagree (malawi.crop_conversion_factors).

An AMBIGUOUS table (two rows sharing a join key) raises ValueError naming the duplicates: de-duplicating is the loader's deliberate act, never an average taken here.

This is NOT Malawi's shelled/unshelled ratio interpolation (EPAR's EPAR_UW_conversionfactors.do:381-390): inferring one condition's factor from the other's is a later, separate layer with its own N threshold (LEARNINGS L9, GH #854). This layer only looks a factor up.

None

Returns:

Type Description
DataFrame

Indexed like crop_production, with

kg_per_unit the resolved factor (NaN for none rows). Deliberately NOT called KgFactor: that name is reserved for the survey's own reported number, and this column is CONSTRUCTED. KgFactorSource one of :data:KG_FACTOR_LAYERS. kg_reported / kg_shipped / kg_survey_median / kg_inferred what each layer offered for that row, whether or not it won. kg_reported and kg_shipped are POST-screen, so a rejected value reads NaN here and kg_reported_rejected / kg_shipped_rejected is True; the raw value is untouched in the input frame's own KgFactor column (and in the caller's own factor table) -- which is what makes the disagreement auditable per unit (groupby('u') on this frame). kg_shipped_source the shipped table's own Source string for the row, where it has one -- Malawi's IHS5 ladder stamps a provenance string on every factor and it is worth carrying through. MISSING wherever no usable shipped factor was found, including where the screen rejected it -- read it with pd.isna, since pandas 3.0 normalises both None and pd.NA to NaN in an object column.

Two tallies ride on .attrs:

kg_factor_sources {layer: n_rows} over the INPUT rows. The five layers of :data:KG_FACTOR_LAYERS partition the frame and sum to len(crop_production) -- shipped included, and therefore present as 0 when no table is passed. Three further keys ride OUTSIDE that partition, since a row they count is still served by one of the five: reported_implausible and shipped_implausible count rows whose offered factor the screen REJECTED, and shipped_matched counts rows the shipped table supplied a finite factor for at all -- a JOIN diagnostic taken before the sentinel, the screen and the rank, and the right number to check when a table looks like it did nothing. POOLED: on a cross-country :class:~lsms_library.feature.Feature frame these counts run over every country at once, so reported: 40000 says nothing about WHICH country reported. Read it as a total, never as coverage; the per-row frame answers the real question with one groupby('country'). kg_factor_disagreement For reported_vs_shipped, reported_vs_survey_median and reported_vs_inferred: {'both': n, 'disagree': k, 'share': k/n or None}, where both counts rows for which BOTH layers produced a usable factor and disagree counts those differing from the reported factor by more than :data:KG_FACTOR_DISAGREEMENT_TOLERANCE in relative terms. This is the audit hook: a large reported_vs_inferred share means the library's factor table and the instrument disagree about what a unit weighs, and the instrument is the one that was there. reported_vs_shipped asks the same question of an externally shipped table, which is precisely the audit GH #852 wants: does the World Bank's conversion table agree with what the enumerator wrote down?

Notes

Only the reported and survey-median layers are validity-filtered (finite, > 0). The inferred layer is used exactly as :func:_kg_factor_series returns it, so that a frame with no KgFactor column reproduces the previous behaviour bit for bit rather than merely closely. _get_kg_factors already refuses non-finite and non-positive values at every point where it can mint one, so the asymmetry is a guarantee about identity, not a hole.

hdds

hdds(food_acquired, *, groups)

Household Dietary Diversity Score (FAO HDDS; EPAR number_foodgroup).

METHODOLOGY transform. Counts how many of the twelve FAO food groups (:data:HDDS_GROUPS) the household acquired ANY of during the food module's recall window.

Parameters:

Name Type Description Default
food_acquired DataFrame

food_acquired item feature, grain (t, v, i, j, u, s), with Quantity and/or Expenditure. Rows are counted across ALL acquisition sources s -- purchased, produced, in-kind and other are all consumption.

required
groups Mapping[str, Sequence[str]]

Required. {group name: [j label, ...]} covering exactly the twelve :data:HDDS_GROUPS and every j present in the frame. The j labels are whatever the frame carries -- so if you called food_acquired(labels='Aggregate') they are the country's Aggregate labels, and if you did not they are its fine Preferred Label items (Uganda: 175 of them). See the labels= contract in .claude/skills/add-feature/food-acquired/aggregate-labels/.

required

Returns:

Type Description
DataFrame

One integer HDDS column (0-12) indexed by (t, i).

Notes

The library ships no default mapping, deliberately. Which of a country's j labels is a "vegetable" is a curation over that country's own food_items / harmonize_food table, and LEARNINGS.org L8 is explicit that EPAR's group mapping is "a starting point to re-derive, not copy". A shipped default would be silently wrong for 15 of the 16 food countries. What IS shipped is the group vocabulary and the partition check.

Divergences from the FAO construct, both real:

  1. Recall window. FAO's HDDS (Kennedy, Ballard & Dop 2011) is a 24-hour recall. LSMS food modules are typically 7-day, so an HDDS computed here is systematically HIGHER than the FAO number and is not comparable to a published FAO HDDS. EPAR's own variable label says so out loud -- "number of food groups individual consumed last week" (W5.do, HOUSEHOLD'S DIET DIVERSITY SCORE). The window is the country's, not this function's; check the country's food module before quoting the score.
  2. Acquisition vs. consumption. food_acquired records what entered the household. Where the country's module is a consumption recall (Malawi's hh_mod_g1, "consumed in the past 7 days", split by source) the two coincide; where it is a pure purchase module, food eaten from own stores registers no row and the score is a lower bound. This function cannot tell the difference -- the analyst must.

A household with rows but no positive acquisition scores 0; a household with no rows at all is absent from the output (never zero-filled).

legacy_area_output

legacy_area_output(country)

Reproduce the retired EthiopiaRHS Country(X).area_output() (t,i) table.

The HH-level wide area_output table (#277) was retired in favour of the item-level crop_production (t, i, j) feature (#438), which subsumes it. This shim re-widens crop_production back to the old (t, i) shape with per-crop {Crop}_kg (production) and {Crop}_ha (area) columns, so callers migrating off area_output() keep working.

Note: crop_production covers MORE waves than the original area_output (R1-R4 + R6 + R7, not just R6/R7), so this shim now returns those extra waves too. New code should use crop_production() directly.

Parameters:

Name Type Description Default
country Country

The Country instance (e.g., ll.Country('EthiopiaRHS')).

required

Returns:

Type Description
DataFrame

DataFrame indexed by (t, i) with {Crop}_kg / {Crop}_ha columns (crop label spaces removed, e.g. WhiteTeff_kg).

legacy_locality

legacy_locality(country)

Reproduce the pre-deprecation output of Country(X).locality().

Returns a DataFrame indexed by (i, t, m) with a single column Parish, where m is the region label and Parish is the parish/cluster identifier (formerly named v in the deprecated interface, renamed in GH #151 to avoid collision with the cluster v used everywhere else in the API).

Implemented by joining sample() and cluster_features() — both first-class tables that carry the same information.

This exists as a compatibility shim for callers migrating off the deprecated locality() method. New code should use sample() and cluster_features() directly.

Parameters:

Name Type Description Default
country Country

The Country instance (e.g., ll.Country('Uganda')).

required

Returns:

Type Description
DataFrame

DataFrame with MultiIndex (i, t, m) and a single column 'Parish'.

livestock_engaged

livestock_engaged(livestock)

Household engaged-in-livestock indicator (WB livestock binary).

MECHANICAL reduction (GAP 4). A household is engaged iff it has any row in the livestock roster: groupby(['t','i']).any().

Parameters:

Name Type Description Default
livestock DataFrame

livestock item feature, grain (t, i, animal).

required

Returns:

Type Description
DataFrame

One boolean Livestock column indexed by (t, i) — True for every HH present in the roster.

Notes

The WB livestock column is 'Yes'/'No'; map this boolean to those strings (.map({True: 'Yes', False: 'No'})) to compare. Because our roster only contains rows for households that own something, every HH in the output is True — the 'No' households are those absent from the roster (present elsewhere in the survey). An analyst recovers the full Yes/No vector by reindexing against the household universe (e.g. sample()) and filling absent HH with False.

livestock_sales_value

livestock_sales_value(livestock, *, price='reported', price_col='ValuePerAnimal', sales_value_col='SalesValue', head_col='HeadSold')

Value of livestock SOLD ALIVE (one term of EPAR's livestock_income).

Named for what it returns, not for EPAR's construct. EPAR's livestock_income is a NET income over seven terms; this is the single value_lvstck_sold term, a GROSS SALES VALUE, and the function name matches its output column Livestock_sales_value so the call site and the column cannot drift apart. See the Notes for the six terms that are missing and why none of them is computable here.

METHODOLOGY transform. HeadSold x price per (t, i, animal), where the price basis is chosen EXPLICITLY by the caller -- because the library's per-head value column is not one price concept but two, and which one a row carries depends on the country and the wave.

Parameters:

Name Type Description Default
livestock DataFrame

livestock item feature, grain (t, i, animal), with a HeadSold column.

required
price (reported, sales_value)

The valuation basis.

  • 'reported' -- multiply HeadSold by price_col (ValuePerAnimal) AS THE COUNTRY CARRIES IT. Read the instrument note below before using this: in most countries that is a reservation price, not a transaction price.
  • 'sales_value' -- do not value anything; return the country's own reported SalesValue (gross value of animals sold in the recall window). Where this column exists it is a REPORTED number and should be preferred to any construction, exactly as crop_production.KgFactor and food_acquired.Quantity_kg are preferred over inferred factors. Raises if the country has no such column.
  • a pd.Series -- a caller-supplied per-animal price, indexed by (t, animal) or by animal alone (e.g. the output of a :func:median_price_valuation-style ladder over sale transactions, which is what EPAR does: Malawi IHS Wave 1/EPAR_UW_Malawi_IHS_W1.do:3320-3346 walks ea -> ta -> district -> region -> country, adopting each cell's weighted median price_per_animal at N >= 10). Every (t, animal) with HeadSold > 0 must be priced or a ValueError names the gaps -- no silent partial valuation.
'reported'
price_col str

Column names, overridable for a country that spells them differently.

'ValuePerAnimal'
sales_value_col str

Column names, overridable for a country that spells them differently.

'ValuePerAnimal'
head_col str

Column names, overridable for a country that spells them differently.

'ValuePerAnimal'

Returns:

Type Description
DataFrame

Indexed by (t, i, animal) with

  • Livestock_sales_value -- value of the head sold, nominal local currency. ADDITIVE over animal: groupby(['t','i']).sum() gives the household total.
  • ValuationSource -- per-row provenance, one of 'reported_per_head' / 'reported_sales_value' / 'supplied_price'.

Rows whose HeadSold or price is missing are DROPPED, not zeroed (the never-zero-fill rule). HeadSold == 0 with a price is a genuine zero and is KEPT.

Notes

What this is a term OF, not a replacement for. EPAR's livestock_income (EPAR_UW_Malawi_IHS_W1.do:5959) is

``value_slaughtered + value_lvstck_sold - value_livestock_purchases
+ (milk + eggs + other products + manure sales)
- (hired labour + fodder + vaccine costs)``

i.e. a NET income. This function computes only the value_lvstck_sold term (:3345). The other six are not computable from our livestock schema and saying so is the point:

  • slaughtered -- no head-slaughtered count in any country's livestock;
  • purchases -- HeadAcquired is a head count with no purchase value attached;
  • products (milk, eggs, manure) and expenses (fodder, water, vaccines, hired labour) -- no such columns anywhere in the schema. FINDINGS_taxonomy.org (LIVESTOCK INCOME row) records both as THEIRS-ONLY at the item layer. GhanaSPS's data_scheme.yml:406-414 notes it holds per-species expense and revenue questions that are deliberately unwired because they have no canonical home.

So the returned number is a GROSS SALES VALUE. Do not label it "livestock income" in published output without the six missing terms.

The price column is two different questions. Per lsms_library/data_info.yml (Columns: livestock: ValuePerAnimal), both variants are per head, but they are not the same economics:

  • RESERVATION price -- "if you would sell one of the [ANIMAL] today, how much would you receive?", reported whether or not anything was sold. Nigeria (s11iq3 W1-W4, s11iq7 W5 -- Nigeria/_/data_scheme.yml:337-343), Malawi (ag_r04 -- Malawi/_/CONTENTS.org:336-344), Uganda 2009-10 / 2010-11 (a6aq6).
  • REALISED average -- "what was, on average, the value of each sold?", defined only where the household sold. Uganda 2011-12 onward (a6aq14b / s6aq14b), which dropped the sell-today question; its non-null rate falls from ~96% to ~21% at that wave and every non-null row from then on is conditional on a sale.

A reservation price is a willingness-to-accept, systematically unequal to the price a household actually got. Valuing sales with it is a modelling choice, which is why price= has no silent default that hides which one you used.

Two countries report the sale value directly and should not be constructed at all: Ethiopia (ls_s8aq60, "total value of sales of [LIVESTOCK] in the last 12 months", Ethiopia/_/data_scheme.yml: 196-206) and Mali (s4aq24 / s8b1q14, "valeur brute des ventes") carry SalesValue and no per-head price. Use price='sales_value' there. GhanaSPS has no REPORTED price basis at all: its value question is a HERD total (HerdValue, "current value of these animals if you sold all of them"), and its revenue questions bundle animals with products, so neither ValuePerAnimal nor SalesValue is declared -- deliberately (GhanaSPS/_/data_scheme.yml:406-414). It DOES declare HeadSold (:465, optional: true), so the only basis that works there is a caller-supplied price= Series; neither 'reported' nor 'sales_value' will resolve.

median_price_valuation

median_price_valuation(item_df, geo_levels, *, value_col='Value_sold', qty_col='Quantity_sold', kg_qty=None, quantity_col='Quantity', item_keys=('j',), threshold=10, weight_col=None, volume_as_mass=True, price_col='_unit_price', out_col='Value')

Value item rows at a geography-ladder median unit price (WB valuation).

METHODOLOGY transform (GAP 7). Reproduces the World Bank LSMS-ISA valuation_median_crops ladder (Reproduction_v2 programs.do:4-103): impute a unit price for every item row from the prices actually observed in sale transactions, taking the median within the smallest geographic cell that clears a minimum-observation threshold, then multiplying that imputed price by each row's physical quantity to get a value.

Method (the choice being encoded — documented for review)
  1. Observed unit price p = value_col / kg_qty per item row, where kg_qty is the sold quantity in kilograms. Rows with a zero or missing price are excluded from the median pool (matching the WB replace crop_price_temp = . if ==0), but still receive an imputed price in step 3.
  2. Median ladder. For each (geo_cell, *item_keys) group, count the usable observed prices n. Walking geo_levels from finest to coarsest, then a final national level, the imputed price for a row is the median observed price of the finest cell whose count ≥ threshold. This is exactly the WB cascade: EA → admin_4 → admin_3 → admin_2 → admin_1 → national, where the cell's median is adopted only if it has ≥10 priced observations and no finer cell already qualified. The national median is the unconditional fallback (the WB replace ... if ten_obs_n==0). With weight_col the cell statistic becomes a weighted median (see that parameter); the ladder, the threshold and the fallback are unchanged.
  3. Valuation. out_col = imputed_price × quantity_col for every row (the WB harvest_value = crop_price * harvest_kg).

Why a ladder of medians rather than each household's own price? The WB construct deliberately values all output — including the home-consumed share that was never sold — at a common local market price, so that two households facing the same market are valued identically regardless of how much each happened to sell. The ≥10-obs threshold trades spatial resolution for a stable median: a cell speaks for itself only when enough sales back it; otherwise it borrows its parent's price.

Parameters:

Name Type Description Default
item_df DataFrame

An item feature carrying, per row, a sale value, a sold quantity, and a physical quantity to value — e.g. crop_production (harvest valuation, drives WB harvest_value_LCU), the seed rows of plot_inputs (seed_value), fertilizer rows (inorganic_fertilizer_value), or plot_labor hired rows (hired_labor_value, where the "price" is a wage and the "quantity" is days). Must carry the geography keys named in geo_levels as index levels or columns.

required
geo_levels sequence of str

Geography keys ordered finest → coarsest (e.g. ['v', 'District', 'Region'] for Uganda — v is the EA/cluster analogue of the WB ea_id). Each must resolve to an index level or a column of item_df. A final unconditional national level is always appended internally, so the caller need not list it. The median within a cell is taken over (geo_level, *item_keys).

required
value_col str

Sale-value column (numerator of the observed unit price).

'Value_sold'
qty_col str

Sold-quantity column, in the row's native unit u. Used only when kg_qty is not supplied: the sold quantity is converted to kg via the shared unit machinery so the unit price is per-kg and comparable across rows reported in different containers.

'Quantity_sold'
kg_qty Series

Pre-computed sold quantity in kilograms, aligned to item_df.index. Supply this when the caller has already converted (or when the price basis is not per-kg — e.g. wages per labor-day, where you pass the days Series and leave quantity_col as the days column). When omitted, qty_col is converted to kg via :func:_kg_factor_series.

None
quantity_col str

The physical quantity each row is valued at in step 3 (harvest kg, seed kg, fertilizer kg, labor days). Must already be in the SAME unit as the price denominator (kg for the default per-kg price); convert upstream (e.g. with :func:harvest_kg's machinery) if needed. Pass a Series via kg_qty-style alignment is not supported here — give a column name.

'Quantity'
item_keys sequence of str

Item identity keys the median is stratified by (the WB cropvar; for seeds the WB adds improved — pass ('j', 'improved')). Resolve to index levels or columns.

('j',)
threshold int

Minimum count of priced observations for a cell's median to be adopted (the WB ten_obs ≥ 10).

10
weight_col str

When given (an index level or a column of item_df), every cell median becomes a weighted median. Sort the cell's observed prices ascending; let W be the total weight and C the cumulative weight. The cell's price is the one at the first row where C >= W/2, EXCEPT where C == W/2 exactly at that row (to within a relative tolerance), in which case it is the mean of that price and the next row's. None (the default) runs the unweighted path unchanged.

The weight is per row -- the household weight repeated on each of its item rows, as EPAR's [aw=weight] is. A quantity-weighted median (EPAR Nigeria W4.do:696, gen weight=qty*weight_pop_rururb) is had by passing a column you derived that way.

Observation counting is unaffected: n stays a COUNT OF ROWS, not a sum of weights, so threshold means the same thing on both paths.

Null and non-positive weights. A priced row whose weight is missing or <= 0 takes no part in the weighted median AND is not counted toward threshold -- a row counts for a cell exactly when it can speak for it. It still receives an imputed price, precisely as a row with a missing price does. A warning names how many rows this is.

Relation to the unweighted path. weight_col is a strict GENERALISATION of it: under equal (or all-equal positive) weights the weighted median reproduces Series.median() EXACTLY, for BOTH odd and even cell counts. The exact-tie exception above is what buys that -- with equal weights it fires precisely on an even count and averages the two central prices, exactly as the unweighted path does; an odd count reaches W/2 strictly inside the middle row and returns it. Pinned in both parities by tests/test_median_price_valuation.py. See :func:_weighted_median.

None
volume_as_mass bool

Forwarded to the kg conversion of qty_col when kg_qty is None.

True
price_col str

Name for the imputed-price column added to the returned frame.

'_unit_price'
out_col str

Name for the valued column (price × quantity).

'Value'

Returns:

Type Description
DataFrame

item_df's index with two added columns: price_col (the imputed ladder median unit price) and out_col (price × quantity_col). One row per input row — the caller groups to whatever HH/plot grain the target WB column lives at (e.g. .groupby(['t','i','plot_id']).sum() for harvest_value_LCU). Rows whose quantity_col is missing get a missing out_col.

Notes

Prior art (convergence, and the deltas). Three teams reached the same selection rule independently -- the WB panel's valuation_median_crops, EPAR (Evans School, UW) Technical Report #335 in both its crop and livestock ladders and again in its consumption repo, and this function: the median price of the finest geographic cell with >= threshold observations. EPAR iterates broad-to-narrow with an unconditional overwrite, which selects the same cell. The verified deltas: - EPAR's medians are WEIGHTED, by its population-raked weight_pop_rururb (Uganda W5.do:482-483, Malawi W1.do:504-505, Nigeria W4.do:695-696); ours is unweighted unless weight_col is passed. We do not rake weights. - EPAR GATES its national rung on the same threshold; ours is an unconditional fallback, so no row is left unvalued. - EPAR's LIVESTOCK ladder is the opposite idiom: narrow-to-broad fill with a preference for the household's OWN observed price, ending unconditional. "EPAR's ladder" is ambiguous -- say which. - EPAR's CONSUMPTION repo gates at obs > 10, i.e. N >= 11, where its Ag repo gates obs > 9, i.e. N >= 10 (our default). See slurm_logs/2026-09-09_epar_curation/LEARNINGS.org L5.

Divergence from WB: their ladder is keyed on survey admin_1..admin_4 codes; we accept whatever geography the caller supplies from our cluster_features / sample (v, District, Region), which are the same nesting at possibly different label granularity. The imputed price therefore matches the WB to the extent the geography nesting and the sold-price pool coincide; the resulting value additionally inherits any unit-conversion coverage caveat of the kg quantities it multiplies (see :func:harvest_kg). *_value columns are LCU; deflation to USD is a separate step the WB does downstream (CPI × Atlas FX) and is out of scope here.

nb_plots

nb_plots(plot_features)

Number of plots per household (WB nb_plots).

MECHANICAL reduction (GAP 6 below-the-line). Counts plot_features rows per household.

Parameters:

Name Type Description Default
plot_features DataFrame

plot_features item feature, grain (t, i, plot_id, ...).

required

Returns:

Type Description
DataFrame

One integer Nb_plots column indexed by (t, i).

Notes

The WB Uganda Household dataset leaves nb_plots empty (the count lives implicitly in the Plot table), so for Uganda this is a sanity-check on plot multiplicity rather than a column match; other WB countries populate nb_plots directly.

nitrogen_kg

nitrogen_kg(plot_inputs, *, nitrogen_content=None, volume_as_mass=True)

Kilograms of nitrogen applied per plot (WB nitrogen_kg).

MECHANICAL reduction (GAP 2). For each fertilizer input row, convert the reported Quantity (native unit u) to kilograms, multiply by the fertilizer's nitrogen share, and sum per plot.

Parameters:

Name Type Description Default
plot_inputs DataFrame

plot_inputs item feature, grain (t, i, plot_id, input, j), with a reported Quantity column and a u column (Uganda stores the input unit as a column u, not an index level).

required
nitrogen_content dict[str, float]

Override the class→N-share map (lowercased input label → kg N per kg product). Defaults to :data:_NITROGEN_CONTENT.

None
volume_as_mass bool

Forwarded to the unit→kg conversion.

True

Returns:

Type Description
DataFrame

One Nitrogen_kg column indexed by (t, i, plot_id). Plots with fertilizer input but no convertible-unit row sum to 0; plots with no fertilizer at all are absent. attrs['nitrogen_input_match'] carries the label-join tally -- matched_rows / unmatched_rows / matched_labels / unmatched_labels (label -> row count, descending).

Notes

Label matching is normalised, and what it cannot resolve is LOUD (2026-09-09). The share is keyed off the input label through :func:_normalise_input_label, which tries the normalised label, then the label with a trailing ' fertilizer' removed, then with one added -- because the corpus writes the same product three ways (Urea in Benin, Urea Fertilizer in Malawi, Phosphate where the key is phosphate fertilizer). Exact-equality matching alone silently zeroed 53,210 Malawi rows (NPK / Urea / CAN / DAP Fertilizer), leaving Other Fertilizer at a nominal 0.0 as the only match and returning 331 plots of exactly 0.0 -- a frame that is non-empty, finite and non-negative, so nothing caught it. A label that still resolves to nothing and names itself a fertilizer input now raises :class:NutrientCoverageWarning; the full tally rides on attrs either way. Matching stays EXACT after normalisation: no substring or fuzzy match, so Organic Fertilizer, D Compound Fertilizer and Ethiopia's NPS stay unmatched and get named rather than being assigned an invented nominal share. Divergence from WB: the WB code keys N-share off the product a survey records (urea 0.46, DAP 0.18, NPK 0.17, …). The UNPS questionnaire (and therefore our plot_inputs.input) records only nutrient class, so we apply nominal class shares (:data:_NITROGEN_CONTENT). Nitrogen_kg will diverge from the WB figure on plots whose product mix differs from the class nominal, and shares the unit-conversion coverage caveat of :func:harvest_kg (only metric-named u rows convert). Pass an explicit nitrogen_content to reproduce a particular country's table.

rcsi

rcsi(food_coping, *, weights=None)

Reduced Coping Strategies Index (WFP rCSI; WB/EPAR rcsi).

MECHANICAL reduction over the food_coping item table: a weighted sum of day-counts across the coping strategies, Σ weight[Strategy] × Days, one score per household-wave.

Parameters:

Name Type Description Default
food_coping DataFrame

food_coping item feature, grain (t, i, Strategy), with an integer Days column (0-7, days in the past 7 the strategy was used).

required
weights dict[str, float]

Strategy label -> weight. Defaults to :data:_RCSI_WFP_WEIGHTS, the standard five-strategy WFP formula (LessPreferred + LimitPortion + ReduceMeals + 3×RestrictAdults + 2×BorrowFood).

The strategy set is a property of the questionnaire, not a universal constant -- do not treat the five-term formula as the definition. Ethiopia and Tanzania field an 8-item battery (the standard five plus LimitVariety / NoFood / WholeDay/WholeDayWithout: Ethiopia/_/ethiopia.py:785-794, Tanzania/_/data_scheme.yml:95-98); Nigeria fields 9 (Nigeria/_/nigeria.py FOOD_COPING_ITEMS, adding SleepHungry / WholeDayNoFood). EPAR's OWN Tanzania rCSI is an eight-term variant with different weights on the extra items (Tanzania NPS/Tanzania NPS Wave 5/EPAR_UW_Tanzania_NPS_W5.do:2617, hh_h02a + hh_h02b + hh_h02c + hh_h02d + 3*hh_h02e + hh_h02f*2 + hh_h02g*4 + hh_h02h*4) -- proof that a wider battery does not collapse to the five-term formula by just ignoring the extra columns. The check is symmetric: a Strategy label present in food_coping but absent from weights raises :class:ValueError naming it (never silently scored on just the terms it recognises), and a weights key that names a label food_coping never carries ALSO raises (never silently scored on fewer terms than the formula claims -- a country fielding only 4 of the 5 standard strategies is equally "not the standard five"). Pass an explicit weights dict naming every label present, and no others, for Ethiopia, Tanzania and Nigeria.

None

Returns:

Type Description
DataFrame

One float rCSI column indexed by (t, i).

Notes

Missing strategy rows are NEVER filled with zero. A household using a strategy zero days is a valid, STORED response (Days=0); food_coping is sparse only where a strategy question went unanswered or its response was out-of-domain and dropped. Evidence:

  • Ethiopia's own wave-builder docstring (Ethiopia/_/ethiopia.py:814-815): "Rows with a missing Days value are dropped (the strategy was not answered for that household)."
  • Nigeria's (Nigeria/_/nigeria.py:985-987): "Rows where the day count is missing are dropped; a household with all-missing items contributes no rows."
  • Malawi's audited row counts (Malawi/_/CONTENTS.org:24-35): rows are exactly 5x each wave's Module-H household count in 2010-11 (12,271 x 5 = 61,355) and short by a handful in the other three waves (6 / 3 / 13 rows) -- the out-of-range day-counts malawi.py coerces to NaN and drops (Malawi/_/malawi.py:1204,1212), never zero-filled.

So a household present for some weighted strategies and absent for another had that strategy UNANSWERED, not zero, and is excluded from the score entirely -- never scored on a truncated sum. Any (t, i) missing so much as one of the strategies named in weights is dropped, not imputed.

rcsi_phase

rcsi_phase(rcsi, *, cutoffs=(3, 18, 42))

WFP rCSI severity phase, in partitioning closed intervals.

Parameters:

Name Type Description Default
rcsi Series or DataFrame

rCSI score(s) -- e.g. the rCSI column returned by :func:rcsi (the frame's single column is used if a DataFrame is passed).

required
cutoffs (int, int, int)

Three ascending cut points splitting the non-negative score line into four phases: [0, c1], (c1, c2], (c2, c3], (c3, inf).

(3, 18, 42)

Returns:

Type Description
Series

Ordered categorical ('Phase 1' .. 'Phase 4'), same index as rcsi.

Notes

Deliberately closed-interval and PARTITIONING -- every real number (hence every integer score) lands in exactly one phase. This avoids a real defect in EPAR's own cutoffs (Malawi IHS/Malawi IHS Wave 1/EPAR_UW_Malawi_IHS_W1.do:4530-4531): phase 2 there is rcsi <= 18 and phase 3 is rcsi > 19 & rcsi <= 42 -- an rcsi of exactly 19 satisfies neither test and falls in NO phase, even though the phase-3 LABEL one line later (:4535) claims the range is "(19 - 42)", i.e. the label and the test contradict each other. Here cutoffs=(3, 18, 42) default gives [0, 3], (3, 18], (18, 42], (42, inf) -- 19 lands in the third phase ((18, 42]), matching the LABEL EPAR intended rather than the gap its test actually produced.

roster_to_characteristics

roster_to_characteristics(df, age_cuts=(4, 9, 14, 19, 31, 51), drop='pid', final_index=['t', 'v', 'i'], mover_sentinel='Mover')

Collapse a household roster into household-level sex × age counts.

Drives the derived household_characteristics table: takes a person-level roster indexed by (at least) ('t', 'v', 'i', 'pid'), buckets each person into a sex_age category, and returns a household-level DataFrame with one integer column per bucket plus a log HSize column (log of household size). Called automatically via :data:lsms_library.country._ROSTER_DERIVED when a user asks Country(name).household_characteristics().

Parameters:

Name Type Description Default
df DataFrame

Household roster with Sex and Age columns (case-insensitive).

required
age_cuts tuple of positive numbers, strictly increasing

Interior breakpoints separating age buckets. See :func:age_intervals for the full semantics; the default (4, 9, 14, 19, 31, 51) produces the historical buckets 00-03, 04-08, 09-13, 14-18, 19-30, 31-50, 51+. Fractional breakpoints (e.g. (0.5, 1, 5)) are allowed and switch the label format from compact 00-03-style to explicit half-open [lo, hi)-style.

(4, 9, 14, 19, 31, 51)
drop str

Index level to drop before aggregation (typically 'pid').

'pid'
final_index list[str]

Final groupby level; defaults to the household key ('t', 'v', 'i') but Country._finalize_result can pass a different tuple when the roster's actual index differs.

['t', 'v', 'i']
mover_sentinel str or None

Value to substitute for NaN entries in any final_index level (in practice always v) before the household groupby. When non-None, mover / split-off households with a NaN cluster code survive as a distinct bucket identifiable by index.get_level_values('v') == mover_sentinel. Pass None to recover the legacy GH #197 behavior where the groupby silently drops these households.

Default flipped from drop-via-NaN to keep-with-sentinel in GH #268. sample() typically carries a valid Region for movers even when v is missing, so dropping them at this stage loses households we /can/ still assign to a market. Callers who want the strict legacy drop pass mover_sentinel=None.

``'Mover'``

Returns:

Type Description
DataFrame

Household-level counts with one column per sex × age bucket and a log HSize column.

Notes

Residence filter. Non-resident members are dropped using whichever residence-duration column the roster carries — MonthsSpent (months present), MonthsAway (→ 12 - value) or WeeksAway (→ 12 - weeks/(52/12)); see CLAUDE.md §"MonthsSpent / MonthsAway / WeeksAway". df here is the country-level concat of every wave, so those columns are unioned across waves: a column supplied by one wave appears (all-NaN) on the others. The resolution is therefore a per-row coalesce across the three sources, and a wave with no usable residence datum at all falls back to counting every roster member — the documented behaviour for a country with no residence column, applied per-wave. Without those two rules an entire wave's household_characteristics silently vanishes while its household_roster is fully populated.

seed_kg

seed_kg(plot_inputs, *, seed_label='Seed', volume_as_mass=True)

Kilograms of seed applied per plot (WB seed_kg).

MECHANICAL reduction (GAP 2). Filters plot_inputs to seed rows, converts the reported Quantity (native unit u) to kilograms, and sums per plot.

Parameters:

Name Type Description Default
plot_inputs DataFrame

plot_inputs item feature (see :func:nitrogen_kg). Seed rows are identified by input == seed_label.

required
seed_label str

The input value marking seed rows (Uganda's harmonize_input Preferred Label).

'Seed'
volume_as_mass bool

Forwarded to the unit→kg conversion.

True

Returns:

Type Description
DataFrame

One Seed_kg column indexed by (t, i, plot_id).

Notes

Coverage: Uganda 2009-10/2010-11 seed rows record no quantity (only purchased-y/n + seed type), so those waves contribute no Seed_kg — matching the WB seed_kg which is also NaN there. Otherwise shares the metric-unit conversion caveat of :func:harvest_kg.

tlu

tlu(livestock, *, tlu_factors=None, head_col='HeadCount')

Tropical Livestock Units owned per household.

MECHANICAL reduction (GAP 4). Σ(HeadCount × species-TLU factor) over the livestock roster. TLU normalises a mixed herd to 250-kg-bovine equivalents (:data:_TLU_FACTORS).

Parameters:

Name Type Description Default
livestock DataFrame

livestock item feature, grain (t, i, animal), with a head- count column.

required
tlu_factors dict[str, float]

Override the species→TLU map (lowercased species label → factor). Defaults to :data:_TLU_FACTORS.

None
head_col str

Head-count column to weight.

'HeadCount'

Returns:

Type Description
DataFrame

One TLU column indexed by (t, i).

Notes

The WB panel ships NO TLU column (its livestock is a bare engaged-y/n binary — see :func:livestock_engaged), so this transform is sanity-checked on magnitudes, not against a WB column. A typical smallholder herd lands at roughly 1-5 TLU; values are bounded below by 0. Species absent from the factor map contribute nothing (and emit no error — callers wanting strict coverage check the species set first).

total_family_labor_days

total_family_labor_days(plot_labor)

Family (household) person-days per HH (WB total_family_labor_days).

MECHANICAL reduction (GAP 3). As :func:total_labor_days but restricted to source == 'family' rows.

total_hired_labor_days

total_hired_labor_days(plot_labor)

Hired (paid) person-days per HH (WB total_hired_labor_days).

MECHANICAL reduction (GAP 3). As :func:total_labor_days but restricted to source == 'hired' rows.

total_labor_days

total_labor_days(plot_labor)

Total person-days of all labor per household (WB total_labor_days).

MECHANICAL reduction (GAP 3). Sums plot_labor.PersonDays over every plot, season, AND source for each household.

Parameters:

Name Type Description Default
plot_labor DataFrame

plot_labor item feature, grain (t, i, plot_id, source, season), with a reported PersonDays column.

required

Returns:

Type Description
DataFrame

One Total_labor_days column indexed by (t, i).

Notes

WB sums to the plot grain (total_labor_days lives on the Plot dataset, one row per plot-season); this returns the household total. Group by plot in the caller (or compare against the WB Plot column summed to HH) to align grains. Coverage caveat: Uganda 2018-19/2019-20 record only hired person-days (no family roster in-repo), so those waves' totals are hired-only — note when comparing to WB, which derives family days from a separate file there.

validate_acquisition_source

validate_acquisition_source(df: DataFrame) -> None

Raise ValueError if the s index level is non-canonical.

No-op when s is not in the index. Called from :meth:lsms_library.country.Country._finalize_result, so it runs on the frame the user actually receives, for every table carrying an s level (food_acquired and the three _FOOD_DERIVED tables built from it).

What this is. A post-condition regression guard, not a bug detector. It passes on 100% of current data: as of 2026-07-12, all 80 (country x food table) frames -- 31.9M rows across the 20 countries with a food_acquired -- are already canonical, with zero NaN. It finds nothing today, and that is the expected steady state. Do not read a green run as evidence that a number is right; read it as evidence that nobody has broken s yet.

Why it is worth its cost anyway. 19 of those 20 countries set s as a bare string literal in a build script (purch['s'] = 'purchased'). A single typo -- 'Purchased', 'purchase', 'own-production' -- would mint a phantom category and silently split rows across the (t, v, i, j, u, s) index, so food_expenditures would quietly drop or double-count a source. That is a silently-wrong-number class of bug, and it is exactly what this guard turns into a loud crash. A crash is a gift.

NaN is a violation, deliberately. The original implementation called .dropna() before comparing, which made it structurally blind to a NaN s -- the dominant real-world non-conformity, since an unmapped source code (EthiopiaRHS/1989 has 677 of them upstream) becomes NaN rather than a bad string. A row whose acquisition source is unknown does not belong on the s axis: it must be mapped to a canonical value, mapped to other, or dropped at the wave level -- explicitly, by the country's build.

See slurm_logs/DESIGN_food_acquired_canonical_2026-05-05.org and GH

169 (which introduced S_VALUES) / GH #537 (which wired this in -- it

had zero call sites from birth despite a docstring claiming otherwise).

yield_kg

yield_kg(crop_production, plot_features, *, area_col='Area', volume_as_mass=True, on='parcel', min_reports=SURVEY_MEDIAN_MIN_REPORTS, shipped_factors=None)

Harvested kilograms per unit plot area (WB yield_kg).

MECHANICAL reduction (GAP 1). Sums :func:harvest_kg and plot area to a common land grain, then divides. Matches the WB yield_kg (harvest summed across crops on a plot/parcel ÷ that land unit's GPS area).

Parameters:

Name Type Description Default
crop_production DataFrame

crop_production item feature (see :func:harvest_kg).

required
plot_features DataFrame

plot_features item feature carrying a plot area column (default 'Area'). Indexed by (t, i, plot_id, ...).

required
area_col str

Name of the area column in plot_features.

'Area'
volume_as_mass bool

Forwarded to :func:harvest_kg.

True
min_reports int

Forwarded to :func:harvest_kg.

:data:`SURVEY_MEDIAN_MIN_REPORTS`
shipped_factors DataFrame

Forwarded to :func:harvest_kg -- an externally shipped kg-per-unit table (see :func:harvest_kg_factors). Without it a caller asking for yields gets no shipped layer, because yield_kg calls harvest_kg internally; that gap was the one concrete cost of the explicit-loader design and this kwarg closes it (.coder/ledger/harvest-kg-shipped-factors.md section 6 Q1/Q3)::

yield_kg(cp, pf, shipped_factors=ethiopia.crop_conversion_factors())
None
on (parcel, plot)

Land grain to join on.

  • 'parcel' (default): reconcile the two plot vocabularies to their common parcel key — crop_production.plot_id is {hhid}-{parcel}-{plot} while plot_features.plot_id is {parcel}_{suffix}, and both encode the same parcel. Harvest is summed over the parcel's crops and plots, area over the parcel's land sub-units (the AGSEC2A/2B split), and divided. This is the WB-faithful grain (their plot_id_merge = hhid-parcel) and on Uganda the parcel key matches on 100% of crop rows.
  • 'plot': join on the literal plot value verbatim. Use only where the two features already share a plot vocabulary; on Uganda the literal ids never match and the result is empty.
'parcel'

Returns:

Type Description
DataFrame

One Yield_kg column indexed by (t, i, parcel) (or (t, i, plot_id) for on='plot') — kilograms per area-unit (the area unit is whatever plot_features.AreaUnit records; Uganda stores hectare-equivalent areas, so Yield_kg is kg/ha).

Notes

Inherits the unit-conversion coverage caveat of :func:harvest_kg (only metric-named u rows convert), so Yield_kg is a lower bound on the WB figure for parcels dominated by non-metric containers. Where the harvest converts and the area key matches, magnitudes track the WB yield_kg distribution.