Skip to content

Visualizations

Three plotting helpers that take a Country, a country name, or the frame the plot is drawn from. All are exported at the package top level (and as ll.visualizations):

import lsms_library as ll

ll.population_pyramid('Uganda', wave='2013-14', ghost=False)        # weighted, unweighted outlined
ll.coordinate_map('Uganda', wave='2013-14', size='weight')           # clusters on a Leaflet basemap
ll.coordinate_map('Uganda', wave='2013-14', size='weight', interactive=False)   # static matplotlib
ll.lorenz_curve('Uganda', wave='2013-14')                            # food spending per person, Gini stated with its basis
ll.lorenz_curve('Uganda', wave=['2009-10', '2013-14', '2019-20'])      # one curve per wave, ordered ramp

matplotlib and folium are ordinary dependencies (since v0.11.0); the interactive map (interactive=True, the default) renders inline in Jupyter or saves as standalone HTML, and interactive=False draws a static matplotlib scatter. folium is imported lazily; matplotlib is loaded at package import anyway (CFEDemands, a core dependency, imports matplotlib.pyplot).

visualizations

Standard charts over harmonized LSMS tables, tuned for a notebook.

Design notes, because the choices are deliberate

Colour never encodes sex. In a population pyramid the mirrored layout already encodes sex unambiguously -- left is one sex, right is the other -- so spending the colour channel on it says nothing the geometry has not already said. Colour is therefore free for the dimension that does carry information: the conditioning variable (by=). Unconditioned, a pyramid is drawn in a single ink. (The conventional blue/pink is also editorially loaded in a way survey demography does not need.)

A chart states the population it represents. Country.population records, per wave, what the survey's own documentation says its sample covers -- and those universes differ: Ethiopia's ESS wave 1 was designed for rural areas and small towns. A pyramid captioned only "Ethiopia 2011-12" invites a reader to treat it as national. The caption is not decoration; it is the difference between a chart and a misleading chart.

The weighting basis is stated, never implied. Weights are normalised to within-wave mean 1 at API time -- but that mean is taken over households, while a roster is one row per person. The weighted person-total is therefore sum(size_h * w_h), which equals the headcount only when weight and household size are uncorrelated, and they are not: GhanaLSS 2016-17 has 59,864 people but a weighted total of 54,421.5, because larger households carry smaller weights. So weighting moves both the shape and the level, and neither is announced by the picture. The subtitle always says which basis produced it.

One wave per pyramid. Pooling waves pools universes (see above), so a multi-wave frame is not silently concatenated: the most recent wave is drawn and the choice is reported.

A Lorenz curve is scale-free, so its basis has to be said. Cumulative shares cancel every unit: currency, deflator, recall period and numeraire= all leave the curve exactly where it was. That is what makes it comparable across waves -- and what lets a quiet choice move the answer with nothing in the picture to betray it. Uganda 2013-14, per person, person-weighted: Gini 0.50 on cash food purchases, 0.34 on all recorded food acquisition, because rural households eat what they grow and cash-only counts none of it (ledger §4; 0.56 household-weighted, 0.48 on household totals unweighted). Whether the poorest half are people or households, whether the survey weights were used, and what became of the households with no recorded purchase each move it again. So lorenz_curve never prints a Gini without naming the measure, the unit, the weighting and the omitted count beside it.

population_pyramid

population_pyramid(data, wave=None, *, weights=True, by=None, bin_width=5, max_age='auto', ghost=None, ax=None, title=None, colors=None)

Draw an age-sex pyramid for one survey wave.

Parameters:

Name Type Description Default
data Country, str, or DataFrame

A country, its name, or a household_roster frame.

required
wave str

Which wave to draw. Defaults to the most recent, with a warning when the frame holds more than one -- their sampled populations may differ.

None
weights bool or str

True uses the survey's sampling weights, falling back to counts with a warning when a country has none. False draws counts. A string names a column to weight by.

True
ghost bool or str

Draw a second weighting as a hairline outline over the filled bars, for comparison. Takes the same values as weights, so the common case is weights=True, ghost=False -- "weighted, with unweighted ghosted". The two series differ in level as well as shape (see the module docstring), and both differences are far easier to read as one overlay than as two charts the eye must hold in memory.

None
by str

Split the pyramid on a category, e.g. by='Rural'. Looked up on the roster, then on sample(). This is what colour encodes.

None
bin_width int

Age-band width in years. Pass 1 to expose age heaping -- LSMS ages pile onto multiples of five, which a 5-year band hides.

5
max_age int or 'auto'

Age at which the closing N+ band starts. "auto" puts it just above the oldest band that would still be visible when drawn: a band thinner than about a pixel adds a row of white space and stretches the vertical scale for nobody. Malawi records ages up to 119 in bands of four people (0.01% of its largest band); drawing to the data maximum would add a dozen empty rows. The test is geometric, not a tuned percentage, so it travels to any country and adapts to the figure size and dpi it is actually drawn at. Pass an integer to fix it.

"auto"
ax matplotlib Axes
None
title optional overrides.
None
colors optional overrides.
None

Returns:

Type Description
Axes

coordinate_map

coordinate_map(data, wave=None, *, size=None, where=None, lat='Latitude', lon='Longitude', label=None, interactive=True, scale='area', colors=None, tiles=None)

Map point coordinates, optionally sizing each marker by a third variable.

The leading case is survey clusters sized by the sampling weight they carry::

coordinate_map('Uganda', wave='2013-14', size='weight')

Given a Country (or its name) the cluster coordinates come from cluster_features and size is summed from sample() over each cluster. Given a DataFrame, lat/lon/size name its columns and nothing is joined.

Restrict the map with where: one region, one stratum, rural clusters only::

coordinate_map('Guinea-Bissau', size='weight', where={'Region': 'bafata'})
coordinate_map('Uganda', wave='2013-14', size='weight',
               where=lambda df: df.Rural == 'Rural')

Parameters:

Name Type Description Default
where dict or callable

{column: value} or {column: [values]} -- every condition must hold -- or a callable taking the assembled frame and returning a boolean mask. For a Country the frame carries the wave's cluster_features columns (Region, Rural, ...) plus the summed size; for a DataFrame, its own columns. A condition that matches nothing raises and lists the values the column holds, so a spelling mismatch is loud. The filter is named in the title, and the no-coordinates count is for the selected clusters only.

None
size str

Column whose value sets marker AREA (not radius -- see :func:_radius_by_area). None draws uniform markers.

None
scale ('area', 'log')

'area' makes marker area proportional to size. 'log' uses log1p(size) instead, which compresses a wide range into a legible one -- Uganda's 197x weight span reads as 5x -- and needs no size floor. It understates differences by construction, so the caption says which was used.

'area'
interactive bool

Render a pan/zoom Leaflet map via folium, which displays inline in Jupyter and can be saved as standalone HTML. False draws a static matplotlib scatter and needs no extra dependency.

True
label str

Column shown in a marker's tooltip (interactive only). Defaults to the cluster id v when present.

None

Returns:

Type Description
``folium.Map`` when ``interactive``, else a matplotlib ``Axes``.
Notes

Coordinates are cluster fixes, not household locations. The published GPS is one point per cluster, stamped onto each of its households; it was never per-household. Mapping households would draw the same point many times and imply a precision the data does not carry. Survey coordinates are also commonly offset before publication to protect respondents, so treat position as approximate.

Clusters without coordinates are counted and reported, never dropped silently: Uganda 2013-14 publishes coordinates for 619 of 706 clusters, so a map that said nothing would omit 12% of the sample without a trace.

lorenz_curve

lorenz_curve(data, wave=None, *, value=None, per='person', weights=True, basis=None, by=None, zeros=None, size=None, ax=None, title=None, colors=None)

Draw the Lorenz curve of household spending for one survey wave.

The cumulative share of spending held by the poorest fraction of the population, on the unit square, against the equality diagonal; the Gini coefficient -- twice the area between them -- is printed on the chart. The curve is scale-free (see the module docstring), so every other choice that moves it is stated in the subtitle: whose spending, per person or per household, weighted how, and what became of households with nothing recorded.

Parameters:

Name Type Description Default
data Country, str, or DataFrame

A country, its name, or a frame. A frame is either an expenditure-shaped table (a j index level and an Expenditure column, summed over every level but (i, t)) or a household-grain frame whose welfare column value names.

required
wave str or list of str

One wave (default: the most recent wave the measure holds, with a warning when it holds several -- Country.waves can list a wave with no food rows), or a list of up to five drawn as separate labelled curves in an ordered ramp of one hue -- most recent in the full ink. A list together with by is an error: colour has one job.

None
value None, str, or pandas.Series

The measure. None is food_expenditures. A string names another table of the country, which must carry an Expenditure column (long shape, summed to household grain); a table without one raises, citing GH #817, rather than summing a wide matrix into a plausible-looking curve. A Series indexed by (i, t) is a user-computed measure -- "food plus non-food" is two API calls and an add, not a kwarg. On the frame path a string names the welfare column.

None
per ('person', 'household')

'person' divides spending by household size and weights each household by size * weight -- the literature's standard, in which each person counts once. 'household' uses household totals weighted by weight. Size is exp(log HSize) from household_characteristics for a Country (resident-filtered, as that table decides), or the size column of a frame.

'person'
weights bool or str

Same contract as :func:population_pyramid: the survey's weights from sample(), falling back to unweighted with a warning; False for unweighted; a column name.

True
basis ('purchased', 'total')

Passed through to food_expenditures: None follows its default (cash purchases only); 'total' is all recorded acquisition value. Meaningful only for the food measure; anywhere else it is an error.

'purchased'
by str or (str, [value, value])

Split into two curves on a column of the frame or of sample() (Rural, strata, Region ...); colour encodes it. A column with more than two values raises and names them; the tuple form picks two of many.

None
zeros ('drop', 'include')

What to do with households in the wave's sample() that have no row in the measure. They are absent, not zero, because food_expenditures drops zero rows (ledger §4) -- on Uganda 2013-14 the 30 such households are true cash zeros with own-production rows. None resolves to 'drop': omit them and say how many. 'include' draws them at zero. Frame path: not accepted -- there is no sample frame to compare against.

'drop'
size str

Frame path only: the household-size column (default 'size').

None
ax matplotlib Axes
None
title optional overrides.
None
colors optional overrides.
None

Returns:

Type Description
Axes
Notes

Counts in the subtitle are of households actually drawn (and, per person, the headcount sum(size) -- never the weighted person total, which is not a count). With several curves the counts are totals over all of them; each curve's Gini sits in its legend entry. Every disclosure ("30 with no recorded purchase, not drawn", "2 without a usable roster, not drawn", "no in-kind value recorded") is a measured count that appears only when it is non-zero.