Working with a Country¶
The Country class is the primary entry point. It
provides access to all survey waves, standardized tables, and panel data for a
single country.
Basic Usage¶
Discovery¶
# Available survey waves
uga.waves
# ['2005-06', '2009-10', '2010-11', '2011-12', '2013-14', '2015-16', '2018-19', '2019-20']
# Available standardized tables
uga.data_scheme
# ['people_last7days', 'cluster_features', 'shocks', 'food_acquired', ...]
Loading Tables¶
Almost every entry in data_scheme is callable as a method:
Two entries are properties, not methods
panel_ids and updated_ids appear in data_scheme but are
@property attributes returning dict, not callables. c.panel_ids()
raises TypeError: 'dict' object is not callable -- use c.panel_ids.
Each returns a pandas DataFrame with a meaningful MultiIndex.
Filtering by Wave¶
Pass waves= to load a subset of survey rounds:
Adding a Market Index¶
For demand-system estimation, add a region-level market identifier:
food = uga.food_expenditures(market='Region')
# Index now includes level `m` derived from cluster_features
Accessing a Single Wave¶
Use bracket notation to get a Wave object:
wave = uga['2019-20']
wave.data_scheme # tables available for this wave
df = wave.household_roster()
Harmonization Pipeline¶
When you call a table method, the library applies these transformations transparently before returning the DataFrame:
- Categorical mappings -- column or index names matching a table in the
country's
categorical_mapping.orgare mapped automatically. A table only maps when it has aPreferred Labelcolumn, and it is keyed on the first column that is neitherPreferred Labelnor a requested label variant -- so write it| Original Label | Preferred Label |(or| Code | Preferred Label |), raw label or code first. A| Code | Label |table with noPreferred Labelis documentation, not a mapping, and changes nothing. - Kinship expansion -- if the wave produces a
Relationshipcolumn, it is decomposed intoGeneration,Distance, andAffinityusing the mapping inlsms_library/categorical_mapping/kinship.yml. Because step 1 runs first, a country#+name: Relationshiptable withPreferred Labelcanonicalises the raw survey wording before the kinship lookup sees it (GH #797; the order was the reverse until then), so one canonical entry inkinship.yml--Spouse-- can serveSpouse (Wife/Husband),Wife/husband, and any other wording a wave uses. The mapping replaces the raw label in the returnedRelationshipcolumn; the cached parquet keeps the raw string, because mappings are a read-path transform (see Caching). One consequence of the order:household_roster(labels={'Relationship': 'X'})feeds the selected label variant into the kinship lookup, not the raw label -- no country curates a multi-labelRelationshiptable today, so this has no live instance. - Canonical spellings -- variant spellings (e.g.
Male->M,Féminin->F) are normalized using the rules indata_info.yml. - Dtype enforcement -- columns are cast to declared types. A country's
data_scheme.ymlis applied first, then the canonicallsms_library/data_info.yml, which wins on conflict (e.g.Age: floatin Albania'sdata_scheme.ymlis overridden toInt64).
Derived Tables¶
Some tables are computed automatically from others:
| Derived Table | Source Table |
|---|---|
food_expenditures |
food_acquired |
food_prices |
food_acquired |
food_quantities |
food_acquired |
household_characteristics |
household_roster |
You call them the same way -- the derivation is transparent.
These are derived tables, computed at read time from a source table. That
is a different thing from a derived value: a number inside a table that the
library constructed from raw survey answers rather than reading off one. Those
are labelled on the row, in a Derivation column carrying the key of the
construction -- see the Derived Values guide. A derived table
does not relabel the rows it inherits; it carries only the
df.attrs['derivations'] summary, pointing at the table where the derivation
was made.
Supported Countries¶
Countries are organized under lsms_library/countries/. To see what's
available: