Transformations¶
Reusable transforms over harmonized tables. The derived tables
(food_expenditures, food_prices, food_quantities,
household_characteristics) are built from these at API time, but the module
is also a library of standalone derived measures — asset indices,
dependency ratios, tropical livestock units, anthropometric z-scores, the World
Bank median-price valuation ladder — that you can call directly.
Check here before writing a new derived measure
Several of these have no call site inside the library yet, so searching for
a usage will not find them. make api-audit lists exactly which.
transformations ¶
A collection of mappings to transform dataframes.
NutrientCoverageWarning ¶
Bases: UserWarning
A fertilizer input label resolved to NO nitrogen share.
:func:nitrogen_kg keys its nutrient share off the plot_inputs.input
label, and a label the map does not recognise contributes NOTHING to the
total -- silently, and with the right shape: the frame comes back
non-empty, finite and non-negative, so every naive assertion passes.
That is the "right shape, no content" failure CLAUDE.md devotes a section
to (null_read_audit), and it happened: before the label
normalisation landed, Malawi's four canonical fertilizer labels (NPK
Fertilizer, Urea Fertilizer, CAN Fertilizer, DAP
Fertilizer -- 53,210 rows) matched none of the map's keys, the ONLY
label that did match was Other Fertilizer at a nominal share of 0.0,
and nitrogen_kg(Malawi) returned 331 rows of exactly 0.0.
Fires for an unmatched label that names itself a fertilizer input --
the corpus's own vocabulary, :data:_FERTILIZER_LABEL_HINTS
(fertilizer / fertiliser / manure / compost, plus the Portuguese
adubo and French engrais) -- and NOT for Seed / Pesticide
/ Herbicide, which are correctly outside a nitrogen map and are the
bulk of the rows (Malawi Seed alone is 112,296). A warning nobody
reads is how the original defect survived, so the firehose form is
deliberately not used. Escalates to a louder message when the whole
returned column is empty or identically zero.
The full per-label tally always rides on
result.attrs['nitrogen_input_match'] whether or not this warns --
that is the coverage check, and it is the twin of
attrs['kg_factor_sources']['shipped_matched'] in
:func:harvest_kg_factors: a JOIN diagnostic, so "did my labels key
correctly?" has an answer that does not depend on a warning firing.
PlotGrainMismatchWarning ¶
Bases: UserWarning
Two plot-level features share no land key, so the join produced nothing.
The corpus runs two plot vocabularies (see :data:_PLOT_LEVELS and
:func:_parcel_from_crop_plot), and a transform that joins across them
on the wrong grain returns an EMPTY frame rather than an error.
Measured: fertilizer_rate(Uganda, on='plot') returned (0, 1)
silently, because the ag modules key plots {hhid}-{parcel}-{plot}
while plot_features keys them {parcel}_{suffix}; the same call
with on='parcel' returns 509 rows.
ShippedFactorWarning ¶
Bases: UserWarning
A shipped_factors table did not join the way its author meant.
Two situations, both of which used to be SILENT and both of which look exactly like a working table from the counts alone:
- the table matched zero rows -- almost always a vocabulary mismatch
on a key (
jas a raw WB crop code against a decoded label, a numeric74.0against'74',tas an int against'2019-20'); - the table is keyed on a level crop_production does not carry, so that level is DROPPED and the factor is applied across it -- a one-region table going national, a one-condition table serving every condition. The multi-row form of this already raised (the surviving keys become ambiguous); the single-row form did not, so the more careless loader got the weaker signal. This makes both loud.
Its own class so it can be filtered or promoted to an error
(warnings.simplefilter('error', ShippedFactorWarning)) independently
of the library's other warnings.
StatedVsInferredKgWarning ¶
Bases: UserWarning
The price-ratio inference disagrees with a unit label's OWN stated size.
A label that names its metric content (Sac moyen (50 kg), Gramme,
Centilitre) is served its STATED factor -- the label is survey
information, the inference is an estimate. When the two disagree by more
than :data:STATED_VS_INFERRED_TOLERANCE, the disagreement is evidence
about the unit vocabulary (a broken kg baseline, a label that means two
sizes) and was previously thrown away. This warning names both numbers.
Its own class so it can be filtered or promoted
(warnings.simplefilter('error', StatedVsInferredKgWarning))
independently of the library's other warnings.
UnitLabelCollisionWarning ¶
Bases: UserWarning
Two u spellings that differ only in case got DIFFERENT kg factors.
conversion_to_kgs groups by the raw u label, so Calebasse and
calebasse are two units; _get_kg_factors then lower-cases and
keeps whichever it merges first. The winner is deterministic (the
inference returns a sorted mapping) but ARBITRARY, and the two factors
can differ several-fold.
Measured over the 19 cached food_acquired tables (GH #770 review):
16 collision groups in 3 countries, and in every one the two factors
DIFFER. Five are INERT -- three kg and two litre, keys that
KNOWN_METRIC seeds, so neither inferred value is ever used --
leaving 11 live groups in 3 countries: Burkina Faso 7, Malawi 2,
Mali 2. Burkina Faso's Calebasse 0.35 vs calebasse 1.432 is
4.1x; Mali's Unité 1.2 vs unité 0.286 is 4.2x. Only the live
groups warn.
Panama is NOT in that list, and was the case that surfaced this:
'Value' -> 0.3489 and 'value' -> 2.3405 was a 6.7x
iteration-order accident. Both spellings are currency-denominated, so
GH #770 removes them from the inference and the collision with them --
which is why the corpus census above is taken AFTER that fix and still
finds 11.
This warns rather than merging, because merging would CHANGE the served factors for those countries -- a data change, not a diagnostic. The fix belongs in the country's own unit vocabulary: canonicalise the spelling where the label is minted, exactly as GH #770 did for currency labels.
UnpriceableRowsWarning ¶
Bases: UserWarning
Emitted when food_prices drops a large share of priceable rows.
Subclasses UserWarning so it can be filtered / promoted to an error
(warnings.simplefilter('error', UnpriceableRowsWarning)) independently
of the library's other warnings.
ValuationLadderWarning ¶
Bases: UserWarning
The own-production valuation ladder ran on fewer geographic rungs than
the canonical v -> District -> Region -> national one.
Emitted by the Country.food_expenditures(valuation=...) path when the
country's cluster_features does not carry one of the coarser rungs (or
the frame carries no v at all), naming which rungs were available. A
shorter ladder is not an error -- median_price_valuation's national
fallback is unconditional, so every row still gets a price -- but it means
the imputed price is drawn from a coarser market than the caller may
assume, so it is said out loud rather than inferred from silence.
age_intervals ¶
Bucket ages into half-open intervals for household_characteristics.
age_cuts is a strictly increasing tuple of interior breakpoints,
each strictly positive. Breakpoints may be fractional (e.g. 0.5 to
separate neonates from older infants). The tuple partitions ages into
len(age_cuts) + 1 half-open buckets:
[0, c_0), [c_0, c_1), ..., [c_{n-1}, inf)
The default (4, 9, 14, 19, 31, 51) reproduces the demographic
buckets 00-03, 04-08, ..., 51+ used throughout the library.
Back-compat: earlier releases took (0, 4, 9, ...) with a leading
zero. A leading zero is now stripped with a :class:DeprecationWarning
so legacy callers keep producing identical buckets.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
age
|
array - like
|
Numeric ages. Negative values fall outside the first bucket and become NaN. |
required |
age_cuts
|
tuple of positive numbers, strictly increasing
|
Interior breakpoints between buckets. |
(4, 9, 14, 19, 31, 51)
|
Returns:
| Type | Description |
|---|---|
Categorical
|
Half-open |
anthropometry_zscores ¶
anthropometry_zscores(anthropometry, who_reference, *, weight_col='Weight', height_col='Height', age_col='Age_months', sex_col='Sex', male_value='M', female_value='F', flag=True)
WHO-2006 child anthropometric z-scores (WB haz06/waz06/whz06/bmiz06).
METHODOLOGY transform (GAP 5). Reproduces the Stata zscore06 /
zanthro call the WB applies (e.g. ETH_ESS1.do:1217-1235) using the WHO
2006 Child Growth Standards: from a child's weight, height/length, age and
sex it computes height-for-age (HAZ), weight-for-age (WAZ),
weight-for-height (WHZ) and BMI-for-age (BMIZ) z-scores, plus the
derived stunting (HAZ < −2) and wasting (WHZ < −2) indicators.
Method (the choice being encoded — documented for review)
The WHO standards distribute each measure as a Box-Cox / LMS family
indexed by sex and age (or, for weight-for-height, by sex and height).
Given the reference parameters L (Box-Cox power), M (median) and
S (coefficient of variation) for a child's (sex, age) cell, the
z-score of a measurement y is
z = ((y / M) ** L − 1) / (L · S) (L ≠ 0)
z = ln(y / M) / S (L = 0)
For the WHO 2006 standards, when |z| > 3 the score is recomputed off
the SD at ±3 to tame the heavy Box-Cox tail (the WHO "restricted" formula
that zscore06 implements):
z = 3 + (y − C₃) / (C₃ − C₂) for z > 3, where Cₖ = M·(1+L·S·k)**(1/L)
z = −3 + (y − C₋₃) / (C₋₂ − C₋₃) for z < −3.
This is the WHO-recommended construction and is exactly what the Stata
zscore06 macro does; encoding that method (rather than a plain
normal-approximation z) is the choice being documented here. With
flag=True (default) z-scores outside the WHO biological-plausibility
bounds (:data:_WHO_FLAG_BOUNDS) are set missing, matching the WB's
flagged-out implausible measurements.
Reference data (NOT vendored)
The WHO LMS reference tables are not shipped with the library (they are
a multi-file WHO dataset under WHO's own terms). The caller supplies them
via who_reference — see that parameter. This keeps the transform
additive and honest: it encodes the WHO-2006 method and leaves the
reference lookup table as an explicit, swappable input, exactly as the
Stata macro relies on the externally-installed WHO ado reference.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
anthropometry
|
DataFrame
|
The |
required |
who_reference
|
dict[str, DataFrame] or callable
|
The WHO 2006 LMS parameter tables. Either:
The igrowup tables are obtainable from
https://www.who.int/tools/child-growth-standards/software (the same
reference the Stata |
required |
weight_col
|
str
|
Measurement columns (kg, cm). |
'Weight'
|
height_col
|
str
|
Measurement columns (kg, cm). |
'Weight'
|
age_col
|
str
|
Child age in months. |
'Age_months'
|
sex_col
|
str
|
Child sex column, values |
'Sex'
|
male_value
|
default 'M' / 'F'
|
Sex labels in |
'M'
|
female_value
|
default 'M' / 'F'
|
Sex labels in |
'M'
|
flag
|
bool
|
Apply the WHO biological-plausibility flagging (set implausible z-scores missing). |
True
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
|
Notes
Divergence from WB: identical method to zscore06; numeric agreement
is to the precision of the supplied reference table and the linear
interpolation between its tabulated x points (WHO tables are at
integer months / 0.1 cm — we interpolate linearly in x, as the WHO
macro does). Because the reference is an explicit input, results match
the WB exactly when fed the same WHO igrowup tables. bmiz06 uses
BMI = weight / (height/100) ** 2.
asset_index ¶
First-principal-component asset index (WB ag/hh_asset_index).
METHODOLOGY transform (GAP 8). Reproduces the World Bank
factor d_*, pcf + predict construct (e.g. ETH_ESS1.do:1077-1102):
reshape the long item-level assets ownership into a household ×
asset-type ownership matrix, extract its first principal component,
and return each household's score on that component as a single wealth
index.
Method (the choice being encoded — documented for review)
Stata's factor ..., pcf is principal-component factoring: the factor
loadings are the eigenvectors of the correlation matrix, and
predict returns the standardised first-component score. We reproduce
that exactly with scikit-learn:
- Build a household × asset-type binary matrix
Dof ownership dummies (1 if the HH owns ≥1 of that asset type, else 0). An asset type the HH never reports is a structural 0, matching the WBreshape wide+ implicit-zero behaviour. - Drop asset-type columns with no variation (all-0 or all-1) — a
constant column has zero correlation contribution and Stata's
factorsilently ignores it. - Standardise each column to mean 0 / unit variance, then run PCA;
operating on standardised columns makes PCA factor the correlation
matrix, which is what
pcfdoes. - The index is the projection onto the first principal component
(
predictafterfactor). Its sign is arbitrary (PCA sign is not identified); we orient it so the component correlates non-negatively with the number of assets owned — a higher index means more assets, the conventional reading.
The score is per (t, i) and computed within each wave t
separately (the WB factors each survey round on its own data), so scores
are not comparable in level across waves — only ranks within a wave are.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
assets
|
DataFrame
|
The |
required |
split
|
(hh, ag)
|
Which index the WB builds.
When |
'hh'
|
value_col
|
str
|
Column to read ownership from when no count column is present.
Defaults to trying |
None
|
ag_items
|
collection of str
|
Asset-type labels ( |
None
|
n_components
|
int
|
Number of leading components to return ( |
1
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
One |
Notes
Divergence from WB: scikit-learn's eigen-decomposition and Stata's
pcf agree on the component direction up to sign and numerical
precision, so the index correlates ~1 with the WB column in rank, but
absolute values differ by the arbitrary sign (we fix it via the
asset-count orientation above) and by tiny scaling/standardisation
conventions. Compare via Spearman rank, not equality. Households absent
from the assets roster (own nothing of the relevant class) are absent
from the result; reindex against the HH universe and fill with the wave
minimum if you need them.
conversion_to_kgs ¶
conversion_to_kgs(df, price=['Expenditure'], quantity='Quantity', index=['t', 'm', 'i'], unit_col='u', *, item_col=None, min_reports=FOOD_KG_MIN_REPORTS, min_baseline=FOOD_KG_MIN_BASELINE, min_baseline_tight=FOOD_KG_MIN_BASELINE_TIGHT, tight_tolerance=FOOD_KG_TIGHT_TOLERANCE, baseline_max_spread=FOOD_KG_BASELINE_MAX_SPREAD, _detail=False)
Infer local-unit -> kg conversion factors from price ratios.
For each unit that does not appear in :data:KNOWN_METRIC, this
function computes a factor by assuming the price per kilogram
should be roughly constant across units for the same item/market.
That is: if a "bunch" of item j trades at roughly 2x the unit value
of a kg of item j, the inferred factor is 2 kg per bunch.
The mechanics, in three steps:
- the baseline -- each row's price per kilogram
(
price / Kgs) is medianed over a group, giving a reference price for a kilogram; - step 2 -- each row's price per UNIT (
price / quantity) is medianed over the same group plusu, over rows whose kilograms are not already known; - the factor -- their ratio, medianed over the remaining axes.
Two things about that chain were wrong before GH #850, and both are fixed here.
Step 2 used to median price itself, not price / quantity,
while this docstring described the result as a per-unit price. The
factor it returned was therefore kilograms per transaction ROW, equal to
kilograms per unit only where Quantity == 1 -- between 4.6% and 41.8%
of the corpus's rows. The self-test is the cheap way to see it: asked
what a kilogram weighs, the old chain answered between 0.87 and 2.0
depending on the country, and the deviation tracked the median quantity
on the rows it looked at.
No step carried the ITEM, so one factor per unit label served a whole
country: Malawi's Piece is shared by 133 food items, and the 72 the
data can estimate separately run from 1 g to 1.8 kg. Pass
item_col='j' for the item-keyed estimate; the default None
reproduces the old keys (and the old return type), with the corrected
arithmetic.
Labels in :data:_CURRENCY_DENOMINATED_UNITS (e.g. u='Value') are
dropped from the whole computation before anything is inferred, so they
never appear as a key of the result: their factor is 1/price at
(j, t, m) and a per-label constant cannot represent it (GH #770).
The item arm's grouping, and why it is not index + ['j']
The baseline is keyed (t, item) -- the WAVE levels of index plus
the item, with the household level dropped. Appending j to index
instead would ask one household to have bought item j both in
kilograms and in bunches within one wave, which is precisely why the
household key fails as an item key. The precedent for dropping the
household is on the crop side: :func:_survey_median_factors groups
(country, t, u, condition) and never carries i.
Two floors gate the item arm, and they gate different things:
min_reports is the number of step-2 rows behind a (t, item, u)
estimate (the twin of :data:SURVEY_MEDIAN_MIN_REPORTS), and
min_baseline is the number of kg-known rows behind the (t, item)
price-per-kg reference. The second is the one that binds. Neither
applies to the item_col=None arm, which is ungated exactly as it was
before #850 -- it is the FALLBACK rung, and gating a fallback would
leave rows with nothing.
min_reports is applied PER WAVE, to each (t, item, u) estimate, and
a wave that cannot clear it contributes nothing to the cross-wave median.
It was applied to the support SUMMED over waves until the GH #850 red team
measured it (2026-09-10), which meant one report per wave in five waves
cleared a floor that four reports in a single wave did not. The crop
side's :func:_survey_median_factors gates per wave; so does this now.
A THIRD gate exists and is OFF by default: baseline_max_spread refuses a
(t, item) baseline whose own reports disagree by more than the given
factor, measured as p90/p10 where the cell has 10 or more reports and
max/min below that. The strict rung is otherwise ungated, and
measured on the corpus it admits cells whose reports differ by up to 9e7x
(median max/min 7.5 in Uganda, 178 in Malawi, 500 in Ethiopia). Left off
pending a measured default; see the sweep in
.coder/ledger/850-food-kg-item-axis-impl.md.
WHAT THE ITEM ARM IS AND IS NOT. It is exactly as good as the item's own
kg-labelled rows. Where those are sound it is right (Niger's cowpea comes
back at 1.11 kg per tiya against an answer key of 1.17); where they are
not, it is wrong AND LOCALISED, whereas the u-pooled factor it replaces
was wrong everywhere and diluted. Niger's millet is the worked example:
nine Kg rows in one wave price millet at 2,286 FCFA/kg against a
retail 250-450, and the item arm hands the resulting 0.33 kg/tiya to all
10,107 of that item's tiya rows. Removing a dilution is not the same
thing as removing an error, and a movement is not a correction until
something outside the estimate says so.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Food-acquired frame with |
required |
price
|
list[str]
|
Column(s) interpreted as expenditure for the ratio calculation. |
['Expenditure']
|
quantity
|
str
|
Column interpreted as quantity. |
'Quantity'
|
index
|
list[str]
|
Groupby levels for the per-item/period median step. At derive time
this is |
['t', 'm', 'i']
|
unit_col
|
str
|
Name of the unit index level; renamed to |
'u'
|
item_col
|
str or None
|
Index level naming the item. |
None
|
min_reports
|
int
|
Floor on the step-2 support behind a |
:data:`FOOD_KG_MIN_REPORTS`
|
min_baseline
|
int
|
Floor on the kg-known rows behind a |
:data:`FOOD_KG_MIN_BASELINE`
|
min_baseline_tight
|
The dispersion-gated exception to min_baseline: a baseline with
between min_baseline_tight and |
FOOD_KG_MIN_BASELINE_TIGHT
|
|
tight_tolerance
|
The dispersion-gated exception to min_baseline: a baseline with
between min_baseline_tight and |
FOOD_KG_MIN_BASELINE_TIGHT
|
|
baseline_max_spread
|
float or None
|
When given, a |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, float] or dict[tuple[str, str], float]
|
Mapping of unit label -> inferred kg factor when With |
crop_diversity ¶
Shannon crop-diversity index per household (EPAR sdi).
METHODOLOGY transform. H = -Sigma_c p_c ln p_c over the crops a
household grew, where p_c is crop c's share of whatever
weight says the shares are shares OF.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
crop_production
|
DataFrame
|
|
required |
weight
|
(count, value, area)
|
What the shares are taken over.
|
'count'
|
value_col
|
str
|
Column used when |
'Value_sold'
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
One float |
Raises:
| Type | Description |
|---|---|
NotImplementedError
|
For |
Notes
Sign. EPAR's sdi is Sigma p ln p -- it never negates
(Uganda UNPS Wave 5/EPAR_UW_Uganda_UNPS_W5.do:3603,3617), so its
published index is NON-POSITIVE and equals -H. Ours is the
conventional Shannon H >= 0. Compare as sdi == -Crop_diversity,
or a reader will think one of the two is broken.
Basis. EPAR weights by area PLANTED (area_plan), dropping rows
with zero planted area -- which, its own comment notes, silently
excludes permanent/tree crops unless they are a plot's only crop. That
weight is refused here rather than approximated (see weight='area'
above).
What one "occurrence" is, stated exactly, because it is easy to
misread. The 'count' basis takes shares over distinct
(t, i, crop, <plot level>, season) tuples -- so a crop grown on
three plots DOES contribute three terms, exactly as in EPAR's plot-crop
row basis. What is de-duplicated away is only the u / condition
multiplicity WITHIN one land-season unit (a maize harvest reported in
two conditions is one occurrence, not two). It is NOT an index over
distinct crops: a household growing maize on four land-season units and
beans on one scores H = 0.5004, not ln 2 = 0.6931. Pinned by
tests/test_863_transforms.py::
test_crop_diversity_count_separates_plots_and_seasons.
Not comparable across countries on the 'count' basis. The
occurrence key uses whichever of the land and season levels the
country's crop_production carries -- (plot, season) for Uganda,
(plot,) for Malawi, which has no season level at all. A Uganda
household with an asymmetric crop mix across its two seasons therefore
scores differently from an otherwise identical Malawi household, purely
because one instrument records a season and the other does not. On a
cross-country :class:~lsms_library.feature.Feature frame, compare
within a country, or use weight='value' (whose shares have no land
or season key) and accept its sold-only bias.
dependency_ratio ¶
Household dependency ratio (WB hh_dependency_ratio).
MECHANICAL reduction (below-the-line, GAP_RANKING Area 3). Per
household: dependents / working-age members, where dependents are
members younger than the working-age band's lower bound or older than its
upper bound.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
household_roster
|
DataFrame
|
|
required |
working_age
|
(int, int)
|
Inclusive |
(15, 64)
|
age_col
|
str
|
Age column name (case-insensitive lookup falls back to lowercase). |
'Age'
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
One float |
Notes
Reproduces hh_dependency_ratio (e.g. ETH_ESS1.do:957-961) up to the
age-bucketing precision of Age. age_handler may return a
fractional Age when DOB is available, so the band test uses < /
> on continuous ages, not integer bins.
dummies ¶
From a dataframe df, construct an array of indicator (dummy) variables, with a column for every unique row df[cols]. Note that the list cols can include names of levels of multiindices.
The optional argument =suffix=, if provided as a string, will append suffix to column names of dummy variables. If suffix=True, then the string '_d' will be appended.
farm_size ¶
Total cultivated/owned plot area per household (WB farm_size).
MECHANICAL reduction (GAP 6 below-the-line). Σ plot Area per
household over plot_features.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
plot_features
|
DataFrame
|
|
required |
area_col
|
str
|
Area column to sum. |
'Area'
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
One |
Notes
WB stores farm_size repeated per plot on the Plot dataset; this
returns the HH total once. Reindex/broadcast to plot grain to compare
directly. Plots with NaN/zero area drop from the sum.
fcs ¶
WFP Food Consumption Score (EPAR fcs).
METHODOLOGY transform. Sigma_g weight[g] x min(days consumed in group
g, 7) per household, over the eight WFP food groups.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
food_acquired
|
DataFrame
|
|
required |
groups
|
Mapping[str, Sequence[str]]
|
Required. |
required |
days
|
str
|
Required, and there is deliberately no default. Name of the column giving the number of days in the past 7 that the item was consumed. No country's |
required |
weights
|
Mapping[str, float]
|
Group -> weight. Defaults to :data: |
None
|
cap
|
int
|
Per-group day cap. A group's item-days are summed, then capped: eating rice on 5 days and maize on 4 is 7 staple-days, not 9. |
7
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
One float |
Notes
Follows WFP's Technical Guidance Sheet (2008): sum the days within each
group, cap the group total at 7, weight, sum. EPAR's Tanzania
implementation diverges from that and the divergence is worth naming:
rather than summing cereals and tubers and capping, it blends them
(W5.do:2579-2582) as 7 if either is 7, else their sum if either
is 0, else (max + min(sum, 7)) / 2 -- an averaging rule with no
citation in the sheet. It also leaves its item-group 2 unweighted, so
that group contributes nothing to the score at all. Neither behaviour
is reproduced here.
The score is NOT thresholded into WFP's poor / borderline / acceptable
bands (21 / 35) by this function -- the same separation as
:func:rcsi / :func:rcsi_phase, so an analyst can apply the
oil-and-sugar-adjusted cutoffs (28 / 42) where the country's diet
warrants them without recomputing the score.
Unlike :func:rcsi, a group with no rows for a household scores 0 days
rather than dropping the household: in a 7-day frequency recall, "no
row" is the instrument's own zero. That is only true if the country's
module really is a consumption recall over the whole item list -- see
:func:hdds's divergence note 2.
fertilizer_rate ¶
fertilizer_rate(plot_inputs, plot_features, *, nutrient='N', area_col='Area', volume_as_mass=True, on='parcel')
Fertilizer applied per unit plot area (EPAR "rate of fertilizer application").
MECHANICAL reduction. Divides a per-plot fertilizer mass by that plot's
area, on the same land grain -- the same grain-bridge shape as
:func:yield_kg.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
plot_inputs
|
DataFrame
|
|
required |
plot_features
|
DataFrame
|
|
required |
nutrient
|
(N, product)
|
|
'N'
|
area_col
|
str
|
Plot-area column in |
'Area'
|
volume_as_mass
|
bool
|
Forwarded to the unit->kg conversion. |
True
|
on
|
(parcel, plot)
|
Land grain to join the two features on. The default matches
:func:
|
'parcel'
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
One column -- |
Notes
Divergences from EPAR's rate, both in the denominator and the numerator, and neither is a bug on either side:
- Denominator. EPAR's is hectares planted, itself imputed as
parcel area times a reported planted share ("Plot area size = Parcel
area size x (planted share)", its readme). Ours is the plot area the
instrument reported and the country converted to hectares
(
plot_features.Area,required). A plot only partly planted therefore gets a LOWER rate here than in EPAR's table. - Numerator. EPAR's rate is over fertilizer PRODUCT; the WB
harmonised panel's
nitrogen_kgis over the nutrient. Hencenutrient=, with no default that hides which one you got. - Inherits :func:
nitrogen_kg's unit-conversion coverage caveat (only metric-namedurows convert) and its class-nominal N shares, so on a country recording nutrient class rather than product the rate is a nominal, not a measured, nitrogen figure.
Plots with fertilizer but zero or missing area are dropped (an infinite rate is not a rate).
food_acquired_valued ¶
food_acquired_valued(df, valuation, *, geo=None, threshold=10, value_col='Expenditure', quantity_col='Quantity')
Value food_acquired's unvalued non-purchased rows, with PROVENANCE.
METHODOLOGY transform (GH #585). The item-grain half of
food_expenditures(basis='total', valuation=...), exposed on its own so
the imputation can be AUDITED row by row -- which rung served each row, and
at what unit price. Every row of df gets a row here, on the same index
and in the same order.
This is opt-in and it is never the default, for a measured reason.
Every rung below prices own-consumption at a PURCHASE transaction, and a
purchase sits above the farm gate by the marketing margin. Our five-source
price study measured the gap: in Uganda, own-consumption valued at the
purchase price sits about 23% above what a sale actually fetched
(slurm_logs/price_sources/SYNTHESIS.org:750-753; the transitive
223-cell set, ratio 1.234). That number is Uganda-only -- GhanaLSS
ships no sale-price source, so no comparable figure exists for it. The
UNPS interviewer manual is explicit that column 9 "should be valued at farm
gate/producer price ... excludes any cost transport and marketing services"
(SYNTHESIS.org:732-736), so a consumption aggregate built this way
carries part of a margin the instrument intended to exclude. Deaton &
Zaidi (2002, Guidelines for Constructing Consumption Aggregates, LSMS
WP 135) treat the producer/market choice as the two defensible readings and
warn that the market one inflates the aggregate relative to purchases.
Choose it deliberately; do not reach for it because basis='total'
looked incomplete.
Rungs, in precedence order (ValuationSource records which one served)
reported
The row already carries a non-zero value_col. Never re-valued --
including a value the survey itself imputed for s='inkind'
(lsms_library/data_info.yml:602). Zero counts as missing, exactly
as :func:food_expenditures_from_acquired's own replace(0, nan)
does, so a zero-valued produced row IS a candidate.
own_price
The same household's own purchase unit value for the same
(t, j, u): the sum of that household's purchased value_col
over the sum of its purchased quantity_col, in that wave, for that
item, in that unit. No cross-household information whatever. This is
EPAR's consumption repo's FIRST rung (EthiopiaW5_...do:423,
UgandaW4_...do:693), conditional there too on the consumed unit
matching the purchased one. Note that EPAR's Ag repo does the
opposite -- it keeps the household's own price as a parallel
value_harvest_hh series rather than a rung -- so "EPAR's ladder" is
ambiguous and this docstring says which (LEARNINGS.org L5 item 4).
median_price
The geographic median-price ladder, run by :func:median_price_valuation
on a price pool built only from s == 'purchased' rows
(value_col / quantity_col). The ladder walks geo finest ->
coarsest and takes the median of the finest cell with at least
threshold priced observations, with an unconditional national
fallback. Cells are (geo, t, j, u): t is in the item key
because median_price_valuation has no wave axis of its own, and
without it the national rung would pool currency across a decade.
none
No rung produced a price. Counted, never hidden. Also covers a
purchased row with no recorded outlay, which is a data defect
rather than an unvalued acquisition and is deliberately not a
candidate.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
|
required |
valuation
|
str or sequence of str
|
Rungs to run, IN THE ORDER GIVEN; each fills only rows still unvalued.
A scalar means exactly that rung -- |
required |
geo
|
DataFrame
|
Coarser geographic rungs, indexed by |
None
|
threshold
|
int
|
Minimum priced observations for a geographic cell's median to be
adopted; forwarded to :func: |
10
|
value_col
|
str
|
Column names on df. |
'Expenditure'
|
quantity_col
|
str
|
Column names on df. |
'Expenditure'
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
Indexed like df, with
Two records ride on
|
Notes
The price basis is the row's NATIVE unit, not kilograms. u is an
item key, so a produced row is valued at the purchase price of the very
label it was reported in and unit alignment is exact by construction. A
kg-normalised basis would pool more finely-split labels (Malawi carries
Kilogramme, Kilogram, Kg and kg as four labels in one wave)
but would inherit the open kg-inference defect of GH #850, and it buys
nothing measurable: national-rung coverage at threshold=10 is 94.2%
(Malawi) / 76.5% (Uganda) on the native key, and lower-casing u moves
it 0.00pp in both. A u='Value' row prices at 1 currency-unit per
currency-unit, which is the right answer for it.
Read that 94.2% as COVERAGE, not as accuracy, and read the
value-weighted complement beside it. median_price_valuation's
national rung is an unconditional fallback, so threshold gates the
geographic cells and gates nothing at the top of the ladder. Measured on
Malawi (red-team, 2026-09-09), by share of the imputed MONEY rather than of
the rows: 6.48% of the imputed value is priced off fewer than ten
purchase observations nationwide, and 2.85% off exactly one. The 5.8% of
candidate rows below the gate are not a random 5.8% of the money. The
concrete case: a single 2016-17 purchase sets "Small Animal - Rabbit, Mice,
Etc." (u='Whole') at 100,000 MWK, which is then applied to 195 rows and
contributes 2.12% of Malawi's entire delta. The median pool is 533
observations, so the bulk is well supported -- but a thin (t, j, u)
cell is thin for every household in it, so this tail is systematic rather
than an outlier. It is COUNTED here, never clipped; valuation_price is
returned per row so a caller can gate on it.
Whether the national rung should carry a gate of its own is a follow-up, deliberately not decided here. @ligon has ruled on the parallel food-side inference (GH #850) that a thin baseline is a floor of 5 with a dispersion-gated exception at 3-4; adopting the same rule at this rung is the obvious candidate, and it would change returned numbers, so it belongs in its own change.
Nothing here is cached. food_expenditures is derived at read time
(Country._FOOD_DERIVED) and this runs inside that derivation, so no
parquet ever holds an imputed value.
food_expenditures_from_acquired ¶
Derive food expenditures from food_acquired.
Returns a DataFrame of expenditure per household × item × period ×
acquisition source (when s is present in the input index), summed
over units.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
food_acquired with an |
required |
basis
|
(purchased, total)
|
Which acquisition sources contribute to
|
'purchased'
|
valuation
|
str or sequence of str
|
Opt-in own-production / in-kind valuation (GH #585). A scalar names EXACTLY that rung: |
None
|
geo
|
DataFrame
|
Coarser geographic rungs for the |
None
|
threshold
|
int
|
Minimum priced observations for a geographic cell's median to be
adopted. Forwarded to :func: |
10
|
Notes
Phase 4 of GH #169 preserves the s (acquisition-source) level in the
output. Users who want the legacy collapsed view call
food_expenditures.groupby(level=['t','v','i','j']).sum() explicitly;
basis='purchased' then yields the purchased total, 'total' the
sum across recorded sources.
When the input has no s level (pre-canonical waves) the two bases
coincide — there is no source split to filter on.
With valuation=, the returned frame carries the per-source tallies on
.attrs['valuation_sources'] (plus 'valuation' and
'valuation_geo_levels'), mirroring harvest_kg's
attrs['kg_factor_sources']. The per-row ValuationSource column is
NOT on this frame: the output is summed over u, and one (t, i, j, s)
cell can mix rungs, so a per-row label has nowhere unique to land. Call
:func:food_acquired_valued directly for the item-grain frame that carries
it.
A :class:~lsms_library.feature.Feature call forwards valuation= (it
forwards by signature) and its VALUES are exact -- Malawi sums to the same
426,562,493.98 either way. Its attrs depend on how many countries were
asked for, and the boundary is worth stating precisely because a reader who
tests the general claim on one country would see it contradicted:
- one country -- a
concatof a single frame, nothing to disagree with, so the tallies COME THROUGH; - more than one -- the frames'
attrsdisagree by construction, one record per country, which lands in the{}row of the propagation rule (CLAUDE.md, "Panel ID Transitive Chains"), soattrs['valuation_sources']is ABSENT.
Pooling the tallies across countries is deliberately NOT built: a pooled
median_price: 280623 would say nothing about WHICH country was imputed,
the same objection harvest_kg's docstring makes about its own pooled
counts. Ask per country, or group the item-grain frame.
food_kg_factors ¶
food_kg_factors(df, *, volume_as_mass=True, item_col='j', unit_kg=None, min_reports=FOOD_KG_MIN_REPORTS, min_baseline=FOOD_KG_MIN_BASELINE, min_baseline_tight=FOOD_KG_MIN_BASELINE_TIGHT, tight_tolerance=FOOD_KG_TIGHT_TOLERANCE, baseline_max_spread=FOOD_KG_BASELINE_MAX_SPREAD)
Per-ROW kg-per-unit for a food_acquired frame, with provenance.
The food twin of :func:harvest_kg_factors, and deliberately the same
shape: one row per input row, a resolved kg_per_unit, a
KgFactorSource drawn from :data:FOOD_KG_FACTOR_LAYERS, the
per-layer candidates beside it, and counts in
attrs['kg_factor_sources'] that PARTITION the frame.
Why per-row at all (GH #850): once the inference carries an item axis, a
factor is no longer a property of the unit label, so there is no dict to
return. And the ladder is now worth reporting -- a row served by its own
(j, u) cell and a row served by the u-pooled fallback were
indistinguishable before, and after this change they differ by more than
they did before it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
|
required |
item_col
|
str
|
The item level. Absent from the frame -> the item rungs serve
nothing and every inferred row falls to |
``'j'``
|
volume_as_mass
|
Passed through; see :func: |
True
|
|
min_reports
|
Passed through; see :func: |
True
|
|
min_baseline
|
Passed through; see :func: |
True
|
|
min_baseline_tight
|
Passed through; see :func: |
True
|
|
tight_tolerance
|
Passed through; see :func: |
True
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
Indexed like df, with columns |
Notes
WHAT THIS DOES NOT FIX, stated where a user will meet it. The only
external answer key the corpus has for a food unit is Mali's own
questionnaire, and against it the corrected inference is BETTER and still
WRONG: for a stated 100 / 50 / 25 kg sack it serves a median 50.0 / 34.4 /
14.1 kg per row (0.50x / 0.69x / 0.56x of truth, up from 0.33x / 0.52x /
0.21x). Gramme is now exact, but only because GH #850 taught the
label parser to READ it -- no inference was involved. Mali's
CONTENTS.org advice ("use native units, or u = 'Kg'") therefore
survives this change. A price-ratio inference cannot be argued into
being a conversion table; where a country ships one, use it (the crop
side's shipped layer, GH #852/#854).
kg_per_unit on a survey_kg row is the factor IMPLIED by the
survey's own kilograms, Quantity_kg / Quantity, so that
Quantity * kg_per_unit reproduces the delivered kilograms on every
row rather than on all but those. Where that quotient cannot be formed
(a missing or zero Quantity) the row keeps the ladder's factor and
still reports survey_kg: the kilograms DO come from the survey, and
the factor column is there for the price-per-kg conversion, which would
otherwise lose a row it can serve.
food_prices_from_acquired ¶
Derive food prices from food_acquired.
Returned at the canonical (t, i, j, u, s) grain (v is omitted
and re-joined from sample() by _finalize_result at API time).
Any extra recall level present in the input — notably GhanaLSS's
visit — is aggregated out here via the median Price, mirroring how
food_quantities_from_acquired / food_expenditures_from_acquired
collapse it with .sum() (price is per-unit, not additive, so median
rather than sum). This keeps the three _FOOD_DERIVED transforms at
the same country-level grain so Feature('food_prices') no longer has
to drop a leaked visit level (and silently keep one arbitrary
visit's price) before the cross-country concat — see GH #517.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
food_acquired DataFrame. Must have |
required |
units
|
(kgvalue, kgprice, unitvalue, unitprice)
|
Which Price to compute, varying across two axes (denominator and source):
See |
'kgvalue'
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
Single-column |
Notes
The 'kgvalue' default deliberately departs from the term-of-art
"unit value" common in the literature (e.g. Deaton 1988, 1997),
which usually means Expenditure / Quantity standardized to kg.
The 'kgvalue' / 'unitvalue' naming makes the denominator
explicit at the cost of mild inconsistency with prior usage; consult
the docstring before substituting one for "unit value" in literature
review or reproduction work.
No silent fallback between modes. 'unitprice' returns NaN where
Price is missing rather than falling back to 'unitvalue'; a
caller wanting "best available" combines results explicitly with
their own provenance tracking.
For canonical s-axis input, only s='purchased' rows have a
meaningful Expenditure (the Phase-3 helpers set produced/inkind
Expenditure to NaN). Under 'kgvalue' / 'unitvalue' those
rows become NaN-after-divide and drop out. Under 'kgprice' /
'unitprice' produced/inkind rows survive iff the wave script
populated the survey-reported Price column upstream (Uganda
Phase 3 path). This function does not synthesize a Price.
food_quantities_from_acquired ¶
Derive food quantities from food_acquired.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
food_acquired DataFrame with a |
required |
units
|
(kgs, units)
|
Aggregation basis:
|
'kgs'
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
Single-column |
Notes
The carry rule for unconvertible units in 'kgs' mode is the
Phase-4 design call recorded in
slurm_logs/DESIGN_food_prices_units_kwarg_2026-05-06.org. It
differs from the original implementation, which silently dropped
unconvertible rows from food_quantities.
The output's group-by preserves both u (the per-row unit; 'kg'
for converted rows, native otherwise) and s (acquisition source),
per the GH #169 canonical schema. Pre-canonical waves where s is
absent silently skip the s level.
format_interval ¶
Render a pd.Interval as a human-readable column-name label.
Two styles, selected by the compact flag:
- Compact (default) — matches the historical
household_characteristicscolumn names. Finite integer bounds span-1 apart or wider collapse to"{lo:02d}-{hi-1:02d}"(e.g.[0, 4)→"00-03"); the unbounded-top bucket renders as"{lo:02d}+"(e.g.[51, inf)→"51+"). - Explicit — half-open interval notation
"[lo, hi)"with"{lo}+"for the unbounded top bucket. Used whenever any bound is fractional (spans shorter than a year).
:func:roster_to_characteristics picks compact=True when all
age_cuts are integers and False otherwise, so column names stay
consistent within a single call.
gross_crop_revenue ¶
Reported crop SALE revenue per household (EPAR value_crop_sales).
MECHANICAL reduction. Sums crop_production.Value_sold -- the sale
value the household actually REPORTED -- to the (t, i) grain, or to
(t, i, crop) with by='j'.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
crop_production
|
DataFrame
|
|
required |
by
|
(None, j, crop)
|
|
None
|
value_col
|
str
|
Reported sale-value column to sum. |
'Value_sold'
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
One |
Notes
This is NOT EPAR's "gross value of crop production", and the two must
not be compared as if they were. EPAR's construct
(Uganda UNPS Wave 5/EPAR_UW_Uganda_UNPS_W5.do, GROSS CROP REVENUE
section) values the entire harvest -- including the share never sold,
which for a smallholder panel is most of it -- at a median unit price
imputed from observed sales on a geographic ladder. That is a
METHODOLOGY transform, and the library already has it:
:func:median_price_valuation, whose ladder (finest geographic cell
with >= 10 priced observations, then a national fallback) is the same
one EPAR and the WB both use. Compose the two -- value the harvest with
median_price_valuation, then sum -- to get their quantity. This
function deliberately reports only what the instrument wrote down.
No de-duplication, by design (count, never clip). Where a country
attached a household-crop sale total to more than one physical row, a
plain sum would double-count it. It does not: measured on the warm
corpus (2026-09-09), at Uganda's finest key (t, i, plot_id, j, season)
only 963 of 44,598 groups hold more than one non-zero sale row, and in
just 37 of those do all rows share ONE Value_sold -- the
stamping shape. 926 carry genuinely differing values, which is what
the instrument implies: UNPS question 6 is compound (how much, in what
unit, in what condition) and the sale questions ride the SAME record, so
each condition row carries its own quantity-sold and value-sold
(Uganda/_/CONTENTS.org §"Harvest condition", and the slot-2 column
list s5bq07a_2 / s5bq08_2). Summing across u / condition
/ season is therefore correct, and 895 exact (t, i, crop)
repeats out of 45,592 rows (1.9% of value) are ordinary coincidence --
two plots of one crop sold in equal quantity for equal money.
On Malawi the risk runs the OTHER way: this is a systematic LOWER
BOUND, not a double count. Its sale is a HOUSEHOLD-crop question that
data_scheme.yml records as "attached to single-plot crops only", and
the consequence is measurable: of 87,440 (t, i, crop) groups grown
on ONE plot, 23,206 (26.5%) carry a sale row, against 514 of 19,311
(2.7%) for crops grown on more than one plot. A crop spread over
two or more plots is almost never credited with a sale at all, by
instrument design -- so Malawi's Gross_crop_revenue UNDER-states
sales, and the shortfall is concentrated in exactly the households
farming the most land. Do not read a Malawi total as complete.
A household present in crop_production that reported no sale at all
is ABSENT from the output, never 0 -- sum(min_count=1), the same
never-zero-fill rule as :func:rcsi. Reindex against sample() if
you want the non-sellers as explicit zeros.
harvest_kg ¶
harvest_kg(crop_production, *, volume_as_mass=True, carry_native=False, min_reports=SURVEY_MEDIAN_MIN_REPORTS, shipped_factors=None)
Total harvested kilograms per (t, i, plot_id, j) from crop_production.
MECHANICAL reduction (GAP 1 → WB Plotcrop/Plot harvest_kg).
For each reported harvest row, convert the native-unit Quantity to
kilograms, then sum within each plot-crop.
Each row's kg-per-unit factor is built by a LAYERED procedure, in
precedence order: the row's own reported KgFactor (what the
instrument wrote down, e.g. UNPS a5?q6d); a shipped factor from
an externally supplied conversion table, when the caller passes one as
shipped_factors; the median of reported
factors for the same (u, condition) in the same country-wave, where
at least min_reports rows report one; the library's inferred
factor from the shared unit→kg machinery (:func:_get_kg_factors via
:func:_kg_factor_series — the same factors
:func:food_quantities_from_acquired applies to food_acquired); and
otherwise none, in which case the row contributes nothing. The
survey's own number wins wherever it exists, on the same principle by
which food_acquired's exact per-row Quantity_kg already beats an
inferred factor.
Which layer served each row is DISCLOSED, not assumed: the counts ride
on result.attrs['kg_factor_sources'] and the per-row provenance is
available from :func:harvest_kg_factors, which also reports how often
the survey's factor and the library's inferred one DISAGREE. A frame
with no KgFactor column can only reach the inferred layer, so it
gets exactly the result it got before the column existed.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
crop_production
|
DataFrame
|
The |
required |
volume_as_mass
|
bool
|
Forwarded to :func: |
True
|
min_reports
|
int
|
How many rows of a |
:data:`SURVEY_MEDIAN_MIN_REPORTS`
|
shipped_factors
|
DataFrame
|
An externally shipped kg-per-unit table, forwarded verbatim to
:func: |
None
|
carry_native
|
bool
|
When False (default, matching the WB construct), rows whose |
False
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
One
|
Notes
Divergence from WB: the WB Uganda code applies survey-provided
conversion tables that cover bespoke local containers ("Sack (100 kgs)",
"Basket (Unspecified)", regional "Heap"/"Bunch" sizes). Our shared
factor map only resolves units that name their metric content (e.g.
"Basket (10 kg)", "Nice cup (60g)") plus the hand-coded metric tokens,
so on Uganda it converts ~21% of crop rows through that layer alone
(28 147 of 133 683, measured 2026-09-09). Harvest_kg is therefore a
lower bound on the WB figure for plots dominated by non-metric
containers; magnitudes agree on metric-reported plots.
Closing that gap is a per-country job in three forms, none of them a
change to this transform: extend the country's u-table
(harvest_units); wire the instrument's own conversion factor into
the optional KgFactor column, which the reported and
survey_median layers above then consume; or -- where the survey ships
a conversion TABLE rather than a per-row factor -- load that table and
pass it as shipped_factors. The third is the one Ethiopia and Malawi
need (GH #852, #854): their tables are keyed on crop x unit x region, so
they can never be a per-row column, and the de-duplication and region
resolution they require are decisions that belong at the call site rather
than baked into a parquet. The second reaches rows the
first cannot -- Uganda's 2018-19 season-A harvest side ships a
100%-populated a5aq6d conversion factor and NO harvest-unit code at
all, so its 7 153 rows sit at u='Unknown' and no unit table can ever
convert them.
The survey_median layer REFUSES the missing-unit sentinel, on both
sides. u='Unknown' (:data:U_UNKNOWN) is the absence of a unit, not
a unit, so a "group" of such rows pools containers of unrelated sizes: a
median over them is a number about nothing, and filling other such rows
with it would fabricate weights. A row with no unit recorded therefore
gets kilograms from its OWN reported KgFactor or from nothing at all
-- which is why wiring KgFactor is the only thing that can ever
convert Uganda's 2018-19 season-A rows.
harvest_kg_factors ¶
harvest_kg_factors(crop_production, *, volume_as_mass=True, min_reports=SURVEY_MEDIAN_MIN_REPORTS, shipped_factors=None)
Per-row kg-per-unit factor for crop_production, and its PROVENANCE.
The factor half of :func:harvest_kg, exposed on its own so the layers
can be AUDITED -- most usefully, the survey's own reported weights
against the library's inferred ones for the same unit. Every row of
crop_production gets a row here, in the same order and on the same
index.
Layers, in precedence order (KgFactorSource records which one served
each row)
reported
The row's own KgFactor column -- what the instrument wrote down
(UNPS a5?q6d, "conversion factor into kg"), where it is finite,
> 0 and PLAUSIBLE (:func:_screen_reported_factors: a kilogram unit
must weigh 1, nothing may exceed :data:KG_FACTOR_MAX, and an
integral calendar year on a non-kg unit is the adjacent-column
signature). A rejected value is never clipped or rescaled -- the row
falls through to the layers below exactly as if it had not reported,
and the rejection is counted under reported_implausible.
KgFactor is REPORTED, never constructed: see the canonical schema
note in lsms_library/data_info.yml.
shipped
A factor looked up in shipped_factors -- an EXTERNALLY SHIPPED
conversion table the caller passed in (the World Bank's Ethiopia
Crop_CF_Wave{2..5}, Malawi's IHS5 crop files: GH #852, #854).
Screened by the SAME :func:_screen_reported_factors, with its
rejections counted under shipped_implausible; a rejected shipped
factor is likewise never clipped, and the row falls through.
Ranked below reported because a table is not this row's
enumerator, and above survey_median because it is external
evidence rather than a pooling of the survey's own reports. See
shipped_factors below for the table's shape and the join.
survey_median
The median reported KgFactor of the same (u, condition)
within the same country-wave, where at least min_reports rows
report one; falling back to u alone. The survey's own weights
filling the survey's own gaps -- see :func:_survey_median_factors.
inferred
The shared unit→kg machinery, unchanged: :func:_kg_factor_series
→ :func:_get_kg_factors (hand-coded metric tokens, the
explicit-metric label parser, then price-ratio inference). This is
the ONLY layer that acts on a frame with no KgFactor column, so
such a frame gets exactly the factors it got before this function
existed.
none
No layer produced a usable factor. Counted, not hidden: these rows
contribute nothing to Harvest_kg.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
crop_production
|
DataFrame
|
The |
required |
volume_as_mass
|
bool
|
Forwarded to :func: |
True
|
min_reports
|
int
|
|
:data:`SURVEY_MEDIAN_MIN_REPORTS`
|
shipped_factors
|
DataFrame
|
An externally shipped kg-per-unit table. NOTHING AUTO-DISCOVERS A COUNTRY'S TABLE, and the result is never stored. The intended call shape, once the loaders of GH #852 / #854 exist, is explicit:: A
An AMBIGUOUS table (two rows sharing a join key) raises This is NOT Malawi's shelled/unshelled ratio interpolation (EPAR's
|
None
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
Indexed like crop_production, with
Two tallies ride on
|
Notes
Only the reported and survey-median layers are validity-filtered
(finite, > 0). The inferred layer is used exactly as
:func:_kg_factor_series returns it, so that a frame with no
KgFactor column reproduces the previous behaviour bit for bit rather
than merely closely. _get_kg_factors already refuses non-finite and
non-positive values at every point where it can mint one, so the
asymmetry is a guarantee about identity, not a hole.
hdds ¶
Household Dietary Diversity Score (FAO HDDS; EPAR number_foodgroup).
METHODOLOGY transform. Counts how many of the twelve FAO food groups
(:data:HDDS_GROUPS) the household acquired ANY of during the food
module's recall window.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
food_acquired
|
DataFrame
|
|
required |
groups
|
Mapping[str, Sequence[str]]
|
Required. |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
One integer |
Notes
The library ships no default mapping, deliberately. Which of a
country's j labels is a "vegetable" is a curation over that
country's own food_items / harmonize_food table, and
LEARNINGS.org L8 is explicit that EPAR's group mapping is "a
starting point to re-derive, not copy". A shipped default would be
silently wrong for 15 of the 16 food countries. What IS shipped is the
group vocabulary and the partition check.
Divergences from the FAO construct, both real:
- Recall window. FAO's HDDS (Kennedy, Ballard & Dop 2011) is a
24-hour recall. LSMS food modules are typically 7-day, so
an HDDS computed here is systematically HIGHER than the FAO number
and is not comparable to a published FAO HDDS. EPAR's own variable
label says so out loud -- "number of food groups individual consumed
last week" (
W5.do, HOUSEHOLD'S DIET DIVERSITY SCORE). The window is the country's, not this function's; check the country's food module before quoting the score. - Acquisition vs. consumption.
food_acquiredrecords what entered the household. Where the country's module is a consumption recall (Malawi'shh_mod_g1, "consumed in the past 7 days", split by source) the two coincide; where it is a pure purchase module, food eaten from own stores registers no row and the score is a lower bound. This function cannot tell the difference -- the analyst must.
A household with rows but no positive acquisition scores 0; a household with no rows at all is absent from the output (never zero-filled).
legacy_area_output ¶
Reproduce the retired EthiopiaRHS Country(X).area_output() (t,i) table.
The HH-level wide area_output table (#277) was retired in favour of the
item-level crop_production (t, i, j) feature (#438), which subsumes it.
This shim re-widens crop_production back to the old (t, i) shape with
per-crop {Crop}_kg (production) and {Crop}_ha (area) columns, so
callers migrating off area_output() keep working.
Note: crop_production covers MORE waves than the original area_output
(R1-R4 + R6 + R7, not just R6/R7), so this shim now returns those extra
waves too. New code should use crop_production() directly.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
country
|
Country
|
The Country instance (e.g., |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame indexed by (t, i) with |
legacy_locality ¶
Reproduce the pre-deprecation output of Country(X).locality().
Returns a DataFrame indexed by (i, t, m) with a single column
Parish, where m is the region label and Parish is the
parish/cluster identifier (formerly named v in the deprecated
interface, renamed in GH #151 to avoid collision with the cluster
v used everywhere else in the API).
Implemented by joining sample() and cluster_features() — both first-class tables that carry the same information.
This exists as a compatibility shim for callers migrating off the deprecated locality() method. New code should use sample() and cluster_features() directly.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
country
|
Country
|
The Country instance (e.g., |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame with MultiIndex (i, t, m) and a single column 'Parish'. |
livestock_engaged ¶
Household engaged-in-livestock indicator (WB livestock binary).
MECHANICAL reduction (GAP 4). A household is engaged iff it has any row
in the livestock roster: groupby(['t','i']).any().
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
livestock
|
DataFrame
|
|
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
One boolean |
Notes
The WB livestock column is 'Yes'/'No'; map this boolean to
those strings (.map({True: 'Yes', False: 'No'})) to compare. Because
our roster only contains rows for households that own something, every
HH in the output is True — the 'No' households are those absent
from the roster (present elsewhere in the survey). An analyst recovers
the full Yes/No vector by reindexing against the household universe
(e.g. sample()) and filling absent HH with False.
livestock_sales_value ¶
livestock_sales_value(livestock, *, price='reported', price_col='ValuePerAnimal', sales_value_col='SalesValue', head_col='HeadSold')
Value of livestock SOLD ALIVE (one term of EPAR's livestock_income).
Named for what it returns, not for EPAR's construct. EPAR's
livestock_income is a NET income over seven terms; this is the
single value_lvstck_sold term, a GROSS SALES VALUE, and the function
name matches its output column Livestock_sales_value so the call
site and the column cannot drift apart. See the Notes for the six terms
that are missing and why none of them is computable here.
METHODOLOGY transform. HeadSold x price per (t, i, animal),
where the price basis is chosen EXPLICITLY by the caller -- because the
library's per-head value column is not one price concept but two, and
which one a row carries depends on the country and the wave.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
livestock
|
DataFrame
|
|
required |
price
|
(reported, sales_value)
|
The valuation basis.
|
'reported'
|
price_col
|
str
|
Column names, overridable for a country that spells them differently. |
'ValuePerAnimal'
|
sales_value_col
|
str
|
Column names, overridable for a country that spells them differently. |
'ValuePerAnimal'
|
head_col
|
str
|
Column names, overridable for a country that spells them differently. |
'ValuePerAnimal'
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
Indexed by
Rows whose |
Notes
What this is a term OF, not a replacement for. EPAR's
livestock_income (EPAR_UW_Malawi_IHS_W1.do:5959) is
``value_slaughtered + value_lvstck_sold - value_livestock_purchases
+ (milk + eggs + other products + manure sales)
- (hired labour + fodder + vaccine costs)``
i.e. a NET income. This function computes only the
value_lvstck_sold term (:3345). The other six are not
computable from our livestock schema and saying so is the point:
- slaughtered -- no head-slaughtered count in any country's
livestock; - purchases --
HeadAcquiredis a head count with no purchase value attached; - products (milk, eggs, manure) and expenses (fodder, water,
vaccines, hired labour) -- no such columns anywhere in the schema.
FINDINGS_taxonomy.org(LIVESTOCK INCOME row) records both as THEIRS-ONLY at the item layer. GhanaSPS'sdata_scheme.yml:406-414notes it holds per-species expense and revenue questions that are deliberately unwired because they have no canonical home.
So the returned number is a GROSS SALES VALUE. Do not label it "livestock income" in published output without the six missing terms.
The price column is two different questions. Per
lsms_library/data_info.yml (Columns: livestock: ValuePerAnimal),
both variants are per head, but they are not the same economics:
- RESERVATION price -- "if you would sell one of the [ANIMAL] today,
how much would you receive?", reported whether or not anything was
sold. Nigeria (
s11iq3W1-W4,s11iq7W5 --Nigeria/_/data_scheme.yml:337-343), Malawi (ag_r04--Malawi/_/CONTENTS.org:336-344), Uganda 2009-10 / 2010-11 (a6aq6). - REALISED average -- "what was, on average, the value of each sold?",
defined only where the household sold. Uganda 2011-12 onward
(
a6aq14b/s6aq14b), which dropped the sell-today question; its non-null rate falls from ~96% to ~21% at that wave and every non-null row from then on is conditional on a sale.
A reservation price is a willingness-to-accept, systematically
unequal to the price a household actually got. Valuing sales with it
is a modelling choice, which is why price= has no silent default
that hides which one you used.
Two countries report the sale value directly and should not be
constructed at all: Ethiopia (ls_s8aq60, "total value of sales
of [LIVESTOCK] in the last 12 months", Ethiopia/_/data_scheme.yml:
196-206) and Mali (s4aq24 / s8b1q14, "valeur brute des
ventes") carry SalesValue and no per-head price. Use
price='sales_value' there. GhanaSPS has no REPORTED price basis
at all: its value question is a HERD total (HerdValue, "current
value of these animals if you sold all of them"), and its revenue
questions bundle animals with products, so neither ValuePerAnimal
nor SalesValue is declared -- deliberately
(GhanaSPS/_/data_scheme.yml:406-414). It DOES declare HeadSold
(:465, optional: true), so the only basis that works there is a
caller-supplied price= Series; neither 'reported' nor
'sales_value' will resolve.
median_price_valuation ¶
median_price_valuation(item_df, geo_levels, *, value_col='Value_sold', qty_col='Quantity_sold', kg_qty=None, quantity_col='Quantity', item_keys=('j',), threshold=10, weight_col=None, volume_as_mass=True, price_col='_unit_price', out_col='Value')
Value item rows at a geography-ladder median unit price (WB valuation).
METHODOLOGY transform (GAP 7). Reproduces the World Bank LSMS-ISA
valuation_median_crops ladder (Reproduction_v2 programs.do:4-103):
impute a unit price for every item row from the prices actually observed
in sale transactions, taking the median within the smallest geographic
cell that clears a minimum-observation threshold, then multiplying that
imputed price by each row's physical quantity to get a value.
Method (the choice being encoded — documented for review)
- Observed unit price
p = value_col / kg_qtyper item row, wherekg_qtyis the sold quantity in kilograms. Rows with a zero or missing price are excluded from the median pool (matching the WBreplace crop_price_temp = . if ==0), but still receive an imputed price in step 3. - Median ladder. For each
(geo_cell, *item_keys)group, count the usable observed pricesn. Walkinggeo_levelsfrom finest to coarsest, then a final national level, the imputed price for a row is the median observed price of the finest cell whose count ≥ threshold. This is exactly the WB cascade: EA → admin_4 → admin_3 → admin_2 → admin_1 → national, where the cell's median is adopted only if it has ≥10 priced observations and no finer cell already qualified. The national median is the unconditional fallback (the WBreplace ... if ten_obs_n==0). Withweight_colthe cell statistic becomes a weighted median (see that parameter); the ladder, the threshold and the fallback are unchanged. - Valuation.
out_col = imputed_price × quantity_colfor every row (the WBharvest_value = crop_price * harvest_kg).
Why a ladder of medians rather than each household's own price? The WB construct deliberately values all output — including the home-consumed share that was never sold — at a common local market price, so that two households facing the same market are valued identically regardless of how much each happened to sell. The ≥10-obs threshold trades spatial resolution for a stable median: a cell speaks for itself only when enough sales back it; otherwise it borrows its parent's price.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
item_df
|
DataFrame
|
An item feature carrying, per row, a sale value, a sold quantity, and
a physical quantity to value — e.g. |
required |
geo_levels
|
sequence of str
|
Geography keys ordered finest → coarsest (e.g.
|
required |
value_col
|
str
|
Sale-value column (numerator of the observed unit price). |
'Value_sold'
|
qty_col
|
str
|
Sold-quantity column, in the row's native unit |
'Quantity_sold'
|
kg_qty
|
Series
|
Pre-computed sold quantity in kilograms, aligned to |
None
|
quantity_col
|
str
|
The physical quantity each row is valued at in step 3 (harvest kg,
seed kg, fertilizer kg, labor days). Must already be in the SAME unit
as the price denominator (kg for the default per-kg price); convert
upstream (e.g. with :func: |
'Quantity'
|
item_keys
|
sequence of str
|
Item identity keys the median is stratified by (the WB |
('j',)
|
threshold
|
int
|
Minimum count of priced observations for a cell's median to be
adopted (the WB |
10
|
weight_col
|
str
|
When given (an index level or a column of The weight is per row -- the household weight repeated on each of its
item rows, as EPAR's Observation counting is unaffected: Null and non-positive weights. A priced row whose weight is missing
or Relation to the unweighted path. |
None
|
volume_as_mass
|
bool
|
Forwarded to the kg conversion of |
True
|
price_col
|
str
|
Name for the imputed-price column added to the returned frame. |
'_unit_price'
|
out_col
|
str
|
Name for the valued column ( |
'Value'
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
|
Notes
Prior art (convergence, and the deltas). Three teams reached the same
selection rule independently -- the WB panel's valuation_median_crops,
EPAR (Evans School, UW) Technical Report #335 in both its crop and
livestock ladders and again in its consumption repo, and this function:
the median price of the finest geographic cell with >= threshold
observations. EPAR iterates broad-to-narrow with an unconditional
overwrite, which selects the same cell. The verified deltas:
- EPAR's medians are WEIGHTED, by its population-raked
weight_pop_rururb (Uganda W5.do:482-483, Malawi
W1.do:504-505, Nigeria W4.do:695-696); ours is unweighted
unless weight_col is passed. We do not rake weights.
- EPAR GATES its national rung on the same threshold; ours is an
unconditional fallback, so no row is left unvalued.
- EPAR's LIVESTOCK ladder is the opposite idiom: narrow-to-broad fill
with a preference for the household's OWN observed price, ending
unconditional. "EPAR's ladder" is ambiguous -- say which.
- EPAR's CONSUMPTION repo gates at obs > 10, i.e. N >= 11, where its
Ag repo gates obs > 9, i.e. N >= 10 (our default).
See slurm_logs/2026-09-09_epar_curation/LEARNINGS.org L5.
Divergence from WB: their ladder is keyed on survey admin_1..admin_4
codes; we accept whatever geography the caller supplies from our
cluster_features / sample (v, District, Region), which
are the same nesting at possibly different label granularity. The imputed
price therefore matches the WB to the extent the geography nesting and
the sold-price pool coincide; the resulting value additionally inherits
any unit-conversion coverage caveat of the kg quantities it multiplies
(see :func:harvest_kg). *_value columns are LCU; deflation to USD
is a separate step the WB does downstream (CPI × Atlas FX) and is out of
scope here.
nb_plots ¶
Number of plots per household (WB nb_plots).
MECHANICAL reduction (GAP 6 below-the-line). Counts plot_features
rows per household.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
plot_features
|
DataFrame
|
|
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
One integer |
Notes
The WB Uganda Household dataset leaves nb_plots empty (the count
lives implicitly in the Plot table), so for Uganda this is a
sanity-check on plot multiplicity rather than a column match; other WB
countries populate nb_plots directly.
nitrogen_kg ¶
Kilograms of nitrogen applied per plot (WB nitrogen_kg).
MECHANICAL reduction (GAP 2). For each fertilizer input row, convert
the reported Quantity (native unit u) to kilograms, multiply by
the fertilizer's nitrogen share, and sum per plot.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
plot_inputs
|
DataFrame
|
|
required |
nitrogen_content
|
dict[str, float]
|
Override the class→N-share map (lowercased input label → kg N per kg
product). Defaults to :data: |
None
|
volume_as_mass
|
bool
|
Forwarded to the unit→kg conversion. |
True
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
One |
Notes
Label matching is normalised, and what it cannot resolve is LOUD
(2026-09-09). The share is keyed off the input label through
:func:_normalise_input_label, which tries the normalised label, then
the label with a trailing ' fertilizer' removed, then with one
added -- because the corpus writes the same product three ways
(Urea in Benin, Urea Fertilizer in Malawi, Phosphate where
the key is phosphate fertilizer). Exact-equality matching alone
silently zeroed 53,210 Malawi rows (NPK / Urea / CAN /
DAP Fertilizer), leaving Other Fertilizer at a nominal 0.0 as
the only match and returning 331 plots of exactly 0.0 -- a frame
that is non-empty, finite and non-negative, so nothing caught it.
A label that still resolves to nothing and names itself a fertilizer
input now raises :class:NutrientCoverageWarning; the full tally rides
on attrs either way. Matching stays EXACT after normalisation: no
substring or fuzzy match, so Organic Fertilizer, D Compound
Fertilizer and Ethiopia's NPS stay unmatched and get named rather
than being assigned an invented nominal share.
Divergence from WB: the WB code keys N-share off the product a survey
records (urea 0.46, DAP 0.18, NPK 0.17, …). The UNPS questionnaire (and
therefore our plot_inputs.input) records only nutrient class, so we
apply nominal class shares (:data:_NITROGEN_CONTENT). Nitrogen_kg
will diverge from the WB figure on plots whose product mix differs from
the class nominal, and shares the unit-conversion coverage caveat of
:func:harvest_kg (only metric-named u rows convert). Pass an
explicit nitrogen_content to reproduce a particular country's table.
rcsi ¶
Reduced Coping Strategies Index (WFP rCSI; WB/EPAR rcsi).
MECHANICAL reduction over the food_coping item table: a weighted sum
of day-counts across the coping strategies, Σ weight[Strategy] ×
Days, one score per household-wave.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
food_coping
|
DataFrame
|
|
required |
weights
|
dict[str, float]
|
Strategy label -> weight. Defaults to :data: The strategy set is a property of the questionnaire, not a
universal constant -- do not treat the five-term formula as the
definition. Ethiopia and Tanzania field an 8-item battery (the
standard five plus |
None
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
One float |
Notes
Missing strategy rows are NEVER filled with zero. A household using
a strategy zero days is a valid, STORED response (Days=0);
food_coping is sparse only where a strategy question went
unanswered or its response was out-of-domain and dropped. Evidence:
- Ethiopia's own wave-builder docstring
(
Ethiopia/_/ethiopia.py:814-815): "Rows with a missing Days value are dropped (the strategy was not answered for that household)." - Nigeria's (
Nigeria/_/nigeria.py:985-987): "Rows where the day count is missing are dropped; a household with all-missing items contributes no rows." - Malawi's audited row counts (
Malawi/_/CONTENTS.org:24-35): rows are exactly 5x each wave's Module-H household count in 2010-11 (12,271 x 5 = 61,355) and short by a handful in the other three waves (6 / 3 / 13 rows) -- the out-of-range day-countsmalawi.pycoerces to NaN and drops (Malawi/_/malawi.py:1204,1212), never zero-filled.
So a household present for some weighted strategies and absent for
another had that strategy UNANSWERED, not zero, and is excluded from
the score entirely -- never scored on a truncated sum. Any (t, i)
missing so much as one of the strategies named in weights is
dropped, not imputed.
rcsi_phase ¶
WFP rCSI severity phase, in partitioning closed intervals.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
rcsi
|
Series or DataFrame
|
rCSI score(s) -- e.g. the |
required |
cutoffs
|
(int, int, int)
|
Three ascending cut points splitting the non-negative score line
into four phases: |
(3, 18, 42)
|
Returns:
| Type | Description |
|---|---|
Series
|
Ordered categorical ( |
Notes
Deliberately closed-interval and PARTITIONING -- every real number
(hence every integer score) lands in exactly one phase. This avoids a
real defect in EPAR's own cutoffs
(Malawi IHS/Malawi IHS Wave 1/EPAR_UW_Malawi_IHS_W1.do:4530-4531):
phase 2 there is rcsi <= 18 and phase 3 is
rcsi > 19 & rcsi <= 42 -- an rcsi of exactly 19 satisfies
neither test and falls in NO phase, even though the phase-3 LABEL one
line later (:4535) claims the range is "(19 - 42)", i.e. the label
and the test contradict each other. Here cutoffs=(3, 18, 42)
default gives [0, 3], (3, 18], (18, 42], (42, inf) --
19 lands in the third phase ((18, 42]), matching the LABEL EPAR
intended rather than the gap its test actually produced.
roster_to_characteristics ¶
roster_to_characteristics(df, age_cuts=(4, 9, 14, 19, 31, 51), drop='pid', final_index=['t', 'v', 'i'], mover_sentinel='Mover')
Collapse a household roster into household-level sex × age counts.
Drives the derived household_characteristics table: takes a
person-level roster indexed by (at least) ('t', 'v', 'i', 'pid'),
buckets each person into a sex_age category, and returns a
household-level DataFrame with one integer column per bucket plus a
log HSize column (log of household size). Called automatically
via :data:lsms_library.country._ROSTER_DERIVED when a user asks
Country(name).household_characteristics().
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Household roster with |
required |
age_cuts
|
tuple of positive numbers, strictly increasing
|
Interior breakpoints separating age buckets. See
:func: |
(4, 9, 14, 19, 31, 51)
|
drop
|
str
|
Index level to drop before aggregation (typically |
'pid'
|
final_index
|
list[str]
|
Final groupby level; defaults to the household key |
['t', 'v', 'i']
|
mover_sentinel
|
str or None
|
Value to substitute for NaN entries in any Default flipped from drop-via-NaN to keep-with-sentinel in
GH #268. |
``'Mover'``
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
Household-level counts with one column per sex × age bucket and
a |
Notes
Residence filter. Non-resident members are dropped using whichever
residence-duration column the roster carries — MonthsSpent (months
present), MonthsAway (→ 12 - value) or WeeksAway
(→ 12 - weeks/(52/12)); see CLAUDE.md §"MonthsSpent /
MonthsAway / WeeksAway". df here is the country-level concat of
every wave, so those columns are unioned across waves: a column
supplied by one wave appears (all-NaN) on the others. The resolution
is therefore a per-row coalesce across the three sources, and a
wave with no usable residence datum at all falls back to counting every
roster member — the documented behaviour for a country with no
residence column, applied per-wave. Without those two rules an entire
wave's household_characteristics silently vanishes while its
household_roster is fully populated.
seed_kg ¶
Kilograms of seed applied per plot (WB seed_kg).
MECHANICAL reduction (GAP 2). Filters plot_inputs to seed rows,
converts the reported Quantity (native unit u) to kilograms, and
sums per plot.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
plot_inputs
|
DataFrame
|
|
required |
seed_label
|
str
|
The |
'Seed'
|
volume_as_mass
|
bool
|
Forwarded to the unit→kg conversion. |
True
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
One |
Notes
Coverage: Uganda 2009-10/2010-11 seed rows record no quantity (only
purchased-y/n + seed type), so those waves contribute no Seed_kg —
matching the WB seed_kg which is also NaN there. Otherwise shares
the metric-unit conversion caveat of :func:harvest_kg.
tlu ¶
Tropical Livestock Units owned per household.
MECHANICAL reduction (GAP 4). Σ(HeadCount × species-TLU factor) over the
livestock roster. TLU normalises a mixed herd to 250-kg-bovine
equivalents (:data:_TLU_FACTORS).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
livestock
|
DataFrame
|
|
required |
tlu_factors
|
dict[str, float]
|
Override the species→TLU map (lowercased species label → factor).
Defaults to :data: |
None
|
head_col
|
str
|
Head-count column to weight. |
'HeadCount'
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
One |
Notes
The WB panel ships NO TLU column (its livestock is a bare
engaged-y/n binary — see :func:livestock_engaged), so this transform
is sanity-checked on magnitudes, not against a WB column. A typical
smallholder herd lands at roughly 1-5 TLU; values are bounded below by 0.
Species absent from the factor map contribute nothing (and emit no
error — callers wanting strict coverage check the species set first).
total_family_labor_days ¶
Family (household) person-days per HH (WB total_family_labor_days).
MECHANICAL reduction (GAP 3). As :func:total_labor_days but restricted
to source == 'family' rows.
total_hired_labor_days ¶
Hired (paid) person-days per HH (WB total_hired_labor_days).
MECHANICAL reduction (GAP 3). As :func:total_labor_days but restricted
to source == 'hired' rows.
total_labor_days ¶
Total person-days of all labor per household (WB total_labor_days).
MECHANICAL reduction (GAP 3). Sums plot_labor.PersonDays over every
plot, season, AND source for each household.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
plot_labor
|
DataFrame
|
|
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
One |
Notes
WB sums to the plot grain (total_labor_days lives on the Plot
dataset, one row per plot-season); this returns the household total.
Group by plot in the caller (or compare against the WB Plot column
summed to HH) to align grains. Coverage caveat: Uganda 2018-19/2019-20
record only hired person-days (no family roster in-repo), so those
waves' totals are hired-only — note when comparing to WB, which derives
family days from a separate file there.
validate_acquisition_source ¶
Raise ValueError if the s index level is non-canonical.
No-op when s is not in the index. Called from
:meth:lsms_library.country.Country._finalize_result, so it runs on the
frame the user actually receives, for every table carrying an s level
(food_acquired and the three _FOOD_DERIVED tables built from it).
What this is. A post-condition regression guard, not a bug detector.
It passes on 100% of current data: as of 2026-07-12, all 80 (country x food
table) frames -- 31.9M rows across the 20 countries with a food_acquired
-- are already canonical, with zero NaN. It finds nothing today, and that
is the expected steady state. Do not read a green run as evidence that a
number is right; read it as evidence that nobody has broken s yet.
Why it is worth its cost anyway. 19 of those 20 countries set s as
a bare string literal in a build script (purch['s'] = 'purchased'). A
single typo -- 'Purchased', 'purchase', 'own-production' --
would mint a phantom category and silently split rows across the
(t, v, i, j, u, s) index, so food_expenditures would quietly drop or
double-count a source. That is a silently-wrong-number class of bug, and it
is exactly what this guard turns into a loud crash. A crash is a gift.
NaN is a violation, deliberately. The original implementation called
.dropna() before comparing, which made it structurally blind to a NaN
s -- the dominant real-world non-conformity, since an unmapped source
code (EthiopiaRHS/1989 has 677 of them upstream) becomes NaN rather than a
bad string. A row whose acquisition source is unknown does not belong on
the s axis: it must be mapped to a canonical value, mapped to other,
or dropped at the wave level -- explicitly, by the country's build.
See slurm_logs/DESIGN_food_acquired_canonical_2026-05-05.org and GH
169 (which introduced S_VALUES) / GH #537 (which wired this in -- it¶
had zero call sites from birth despite a docstring claiming otherwise).
yield_kg ¶
yield_kg(crop_production, plot_features, *, area_col='Area', volume_as_mass=True, on='parcel', min_reports=SURVEY_MEDIAN_MIN_REPORTS, shipped_factors=None)
Harvested kilograms per unit plot area (WB yield_kg).
MECHANICAL reduction (GAP 1). Sums :func:harvest_kg and plot area to a
common land grain, then divides. Matches the WB yield_kg (harvest
summed across crops on a plot/parcel ÷ that land unit's GPS area).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
crop_production
|
DataFrame
|
|
required |
plot_features
|
DataFrame
|
|
required |
area_col
|
str
|
Name of the area column in |
'Area'
|
volume_as_mass
|
bool
|
Forwarded to :func: |
True
|
min_reports
|
int
|
Forwarded to :func: |
:data:`SURVEY_MEDIAN_MIN_REPORTS`
|
shipped_factors
|
DataFrame
|
Forwarded to :func: |
None
|
on
|
(parcel, plot)
|
Land grain to join on.
|
'parcel'
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
One |
Notes
Inherits the unit-conversion coverage caveat of :func:harvest_kg
(only metric-named u rows convert), so Yield_kg is a lower
bound on the WB figure for parcels dominated by non-metric containers.
Where the harvest converts and the area key matches, magnitudes track
the WB yield_kg distribution.