Skip to content

Latest commit

 

History

3 Commits

Folders and files

Repository files navigation

Call of the Blocs

Mapping post-Cold War UN General Assembly alignment through overlap-weighted MDS, medoid blocs, and country–resolution co-clustering

R UNGA Matrix MDS PAM Co-clustering Validation License: MIT

Read the final report · Build the vote matrix · Run MDS, PAM, and co-clustering · Inspect the custom linkage code

Original academic project title: Call of the Blocs

A UN General Assembly roll call is both a diplomatic event and an entry in a very large relational table. Viewed across three decades, that table can reveal durable country-to-country alignment; viewed resolution by resolution, it can reveal coalitions that reorganize with the issue. Call of the Blocs models both structures without pretending that an unrecorded vote is an observed disagreement.

The first lens compares countries only on roll calls where both have a recorded vote, weights each comparison by how much shared evidence supports it, embeds the resulting dissimilarities with ratio-SMACOF multidimensional scaling, and discovers broad blocs with PAM. The second lens returns to the full country-by-resolution table and jointly clusters both sides with a categorical latent block model, preserving absence as its own state.

176
countries after participation filters
2,609
post-Cold War roll calls
420,230
recorded Yes, Abstain, and No votes
15,400
country pairs with nonzero overlap

This repository is a compact technical guide to the analysis. The report remains the canonical evidence package with every figure, table, appendix, and academic reference; the public R programs make the full computational path inspectable and runnable.

Contents

The substantive question

How is post-Cold War UNGA voting similarity organized across countries, and what recurring coalitions become visible only when the resolution dimension is retained?

The question has a built-in missing-data problem. A country that has no recorded vote on a resolution supplies neither Yes, Abstain, nor No. Pairwise agreement must therefore be calculated over the intersection of two countries' observed roll calls, and comparisons supported by more shared roll calls should carry more weight.

The project is descriptive and unsupervised. It identifies structure in recorded votes; it does not predict future positions, infer governments' motivations, rank countries normatively, or estimate a causal effect of bloc membership.

One voting relation, two analytical lenses

The primitive object is bipartite: delegations connect to resolutions through categorical vote states. A country projection and a block model answer different questions about the same relation.

graph LR
    subgraph Delegations
        I((Country i))
        J((Country j))
    end

    subgraph Roll_calls
        R1{{Resolution 1}}
        R2{{Resolution 2}}
        R3{{Resolution 3}}
    end

    I -- Yes --> R1
    J -- Yes --> R1
    I -- No --> R2
    J -- Abstain --> R2
    I -- Yes --> R3

    R1 --> S[Shared evidence: agreement]
    R2 --> T[Shared evidence: disagreement]
    R3 -.-> U[Excluded: no recorded vote for country j]

    S --> P[Country projection: dissimilarity and overlap]
    T --> P
    I --> B[Full bipartite table]
    J --> B
    R1 --> B
    R2 --> B
    R3 --> B
    B --> L[Country and resolution latent blocks]

    classDef country fill:#E7F5FA,stroke:#0077A3,color:#12313D,stroke-width:2px
    classDef vote fill:#FFF1D6,stroke:#B66A00,color:#3F2B0A,stroke-width:2px
    class I,J country
    class R1,R2,R3 vote
Loading
Global country lens Resolution-sensitive lens
Collapse shared voting records into one country-by-country dissimilarity matrix. Keep the complete country-by-resolution categorical table.
Weight each pair by its number of jointly observed roll calls. Encode non-recorded participation as a distinct Absent state.
Use MDS for continuous geometry and PAM for broad blocs. Estimate country groups, resolution groups, and block-specific vote probabilities jointly.
Answer: who tends to vote alike across the period? Answer: which coalitions recur for which sets of resolutions?

This split prevents an attractive global map from becoming the only model of alignment.

Data lineage and analytic sample

The scripts load un_votes, un_roll_calls, and un_roll_call_issues from the unvotes R package. That package distributes a cleaned form of the Voeten, Strezhnev, and Bailey UN voting archive, whose research-data record is available through Harvard Dataverse. The UN Digital Library voting-data record provides a separate official reference point for General Assembly voting records; it is not silently substituted for the versioned package input.

The construction is deterministic:

  1. keep recorded Yes, Abstain, and No outcomes;
  2. restrict roll calls to 1990–2019;
  3. retain countries with at least 1,500 recorded votes in that window;
  4. retain roll calls with at least 120 participating countries;
  5. preserve issue labels for post-estimation interpretation, never as clustering inputs.
Stage Countries Roll calls Recorded votes
Post-Cold War window before stability filters 194 2,650 444,195
Final analytic sample 176 2,609 420,230
Final recorded state Count Share of recorded votes
Yes 341,442 81.25%
Abstain 51,012 12.14%
No 27,776 6.61%

The preparation stage writes intermediate .rds objects under ignored data_processed/; no copied source-data binary is committed. See report section 2, physical PDF pages 2–3.

Shared evidence before distance

Let i and j index countries and r index roll calls. For the country-level projection, recorded votes are coded as

$$V_{ir} \in \{1,2,3\} \quad \text{for} \quad \{\mathrm{Yes},\mathrm{Abstain},\mathrm{No}\},$$

with V_ir missing when country i has no recorded vote on resolution r. Define the observation indicator

$$M_{ir} = \mathbf{1}\{V_{ir}\text{ is recorded}\}.$$

The number of jointly observed roll calls is

$$O_{ij} = \sum_{r=1}^{m} M_{ir}M_{jr},$$

and exact agreements over that shared evidence are

$$A_{ij} = \sum_{r=1}^{m} \mathbf{1}\{V_{ir}=V_{jr}\}M_{ir}M_{jr}.$$

The baseline dissimilarity and its reliability weight are

$$\delta_{ij} = 1-\frac{A_{ij}}{O_{ij}}, \qquad w_{ij}=O_{ij}, \qquad \delta_{ii}=w_{ii}=0.$$

Thus δ_ij = 0 means complete exact agreement on the countries' shared recorded votes, while δ_ij = 1 means complete disagreement on that shared set. No country pair has zero overlap in the final sample.

Overlap diagnostic Roll calls
Minimum 842
10th percentile 1,727
Median 2,277
90th percentile 2,541
Maximum 2,609

The overlap weight does not alter δ_ij; it changes how strongly that value influences the embedding. A distance based on 2,500 common votes contributes more evidence than one based on 850.

Weighted ratio-SMACOF MDS

Multidimensional scaling assigns country i a coordinate x_i in p dimensions and compares Euclidean map distances with observed voting dissimilarities:

$$d_{ij}(X)=\lVert x_i-x_j\rVert_2, \qquad x_i \in \mathbb{R}^{p}.$$

Ratio MDS estimates one common scale factor through the origin:

$$\widehat{d}_{ij}=b\delta_{ij}, \qquad \widehat{b}(X) = \frac{\sum_{i\lt j}w_{ij}\delta_{ij}d_{ij}(X)} {\sum_{i\lt j}w_{ij}\delta_{ij}^{2}}.$$

The weighted raw-stress loss is

$$\sigma_{\mathrm{raw}}(X) = \sum_{i\lt j}w_{ij} \bigl(\widehat{b}(X)\delta_{ij}-d_{ij}(X)\bigr)^2,$$

and the reported scale-free Stress-1 diagnostic is

$$\sigma_1(X) = \sqrt{ \frac{ \sum_{i\lt j}w_{ij}\bigl(\widehat{b}(X)\delta_{ij}-d_{ij}(X)\bigr)^2 }{ \sum_{i\lt j}w_{ij}d_{ij}(X)^2 } }.$$

SMACOF minimizes stress through iterative majorization. The implementation fits dimensions 1–4 with six random starts each, retains two dimensions at the empirical elbow, and refits the two-dimensional ratio model from 20 deterministic starts. Each fit allows at most 1,000 majorization iterations with tolerance 1e-6; all random starts derive from seed 1363.

Two-dimensional fit diagnostic Exact result
Stress-1 0.103895
Fitted ratio slope 0.067790
Squared weighted Shepard correlation 0.975341

The last quantity is the squared weighted correlation between fitted disparities and configuration distances. It is a descriptive Shepard-fit measure, not an out-of-sample predictive coefficient of determination.

MDS coordinates are invariant to translation, rotation, and reflection. Neither axis direction nor sign has intrinsic political meaning; only inter-country distances, neighborhoods, and broad regions are interpreted. See report section 3 and Figure 2, physical PDF pages 4–7.

Country blocs: PAM on the observed dissimilarities

PAM, or k-medoids, operates directly on the 176-by-176 dissimilarity matrix. For candidate medoids m_1 through m_K, it minimizes

$$\min_{\{m_1,\ldots,m_K\}} \sum_{i=1}^{n} \min_{1\leq k\leq K}\delta_{i,m_k}.$$

Unlike a centroid, each medoid is an observed country. The analysis fits K = 2, ..., 10 and chooses the largest average silhouette width.

PAM result Exact value
Selected number of blocs 2
Average silhouette width 0.650491
Bloc sizes 124 and 52 countries
Medoids Burkina Faso and Slovakia

Cluster numbers are arbitrary identifiers. The medoids are representative observed voting profiles under the fitted dissimilarity—not political leaders, normative centers, or permanent bloc labels.

Abstention sensitivity

Dissimilarity treatment Selected K Best average silhouette ARI against baseline
Exact-match baseline 2 0.650 1.000
Soft abstention costs 2 0.706 1.000
Recode Abstain as No 2 0.689 0.977

All variants preserve the two-bloc solution. The adjusted Rand index shows that the membership partition is identical under soft costs and changes only slightly under the Abstain-as-No coding.

The ten maximally polarizing labeled roll calls in the report show a 100-percentage-point gap in Yes shares between blocs. Human-rights resolutions dominate this particular extreme set, but issue metadata is attached only after clustering. See report section 4 and Appendix Tables 2 and 4, physical PDF pages 7–12.

Resolution-sensitive structure: categorical co-clustering

A single country distance averages across all resolutions. The second lens restores the complete bipartite table and represents non-recorded participation explicitly:

$$Y_{ir} \in \{\mathrm{Yes},\mathrm{Abstain},\mathrm{No},\mathrm{Absent}\}.$$

Country i belongs to row cluster z_i; resolution r belongs to column cluster u_r. Inside country cluster k and resolution cluster g, the categorical block probabilities are

$$\theta_{kg\ell} = \Pr(Y_{ir}=\ell \mid z_i=k, u_r=g), \qquad \sum_{\ell=1}^{4}\theta_{kg\ell}=1.$$

The model grid evaluates all nine combinations with two, three, or four country clusters and two, three, or four resolution clusters. Every fit receives a deterministic model-specific seed. Integrated Completed Likelihood selects the most supported partition in the tested grid.

Co-clustering result Exact value
Selected country clusters 4
Selected resolution clusters 4
Selected ICL -260,386.5
Human-rights share in labeled Resolution Cluster 2 86.8%

The fitted blocks show that Yes, Abstain, No, and Absent probabilities vary across resolution groups. Other discovered column clusters combine human rights, arms control and disarmament, the Palestinian conflict, colonialism, and development in different proportions. These labels interpret the model after estimation; they do not supervise it.

The result is deliberately richer than the global PAM partition: strong broad alignment coexists with issue-dependent coalition and participation patterns. See report sections 3–4 and Figure 5, physical PDF pages 6–10.

Custom agglomerative clustering: a separate algorithm check

The repository also implements hierarchical agglomerative clustering from first principles on a transparent ten-point dissimilarity matrix. At each step it merges the closest active clusters and updates their distance using one of three rules:

$$\begin{aligned} d_{\mathrm{single}}(A,B) &= \min_{i\in A,\,j\in B} d_{ij}, \\\ d_{\mathrm{complete}}(A,B) &= \max_{i\in A,\,j\in B} d_{ij}, \\\ d_{\mathrm{average}}(A,B) &= \frac{1}{|A||B|} \sum_{i\in A}\sum_{j\in B}d_{ij}. \end{aligned}$$

The code maintains the active distance matrix, breaks exact ties deterministically, records the full partition at every merge, and validates every possible dendrogram cut against base R's hclust().

Benchmark diagnostic Single Complete Average
Maximum absolute merge-height gap 0 0 0
Minimum ARI across all tested cuts 1.000 1.000 1.000
Pairwise partition agreement 100% 100% 100%
Cophenetic correlation 0.9954866 0.9955481 0.9956034

This is an algorithm-validation exercise, separate from the substantive UNGA pipeline. The ten-point matrix makes each merge and encoder row inspectable. The full exercise begins at physical PDF page 33.

Results ledger

Layer Estimand or decision Result What it establishes
Pairwise evidence Country overlap 842–2,609 roll calls Every distance is based on shared recorded votes; support varies materially.
Continuous geometry Two-dimensional ratio-SMACOF Stress-1 0.103895 A compact map retains the major pairwise structure.
Broad grouping Silhouette-selected PAM K = 2, silhouette 0.650491 A strong global two-region partition is present.
Robustness Abstention alternatives ARI 0.977–1.000 The broad split is not an artifact of one abstention coding.
Relational grouping Categorical latent blocks 4 × 4, ICL -260,386.5 Coalition and absence profiles vary by resolution group.
Algorithm verification Manual linkage versus hclust() Exact partitions and merge heights The custom clustering implementation reproduces the benchmark.

The empirical conclusion is layered: a strong global two-bloc geometry exists, but it does not exhaust the resolution-specific structure of UNGA voting.

Artifact map

Question or artifact Canonical location
Research question and design Report section 1, physical PDF page 2
Data source, coding, and filters Report section 2, physical PDF pages 2–3
Dissimilarity, MDS, PAM, and latent-block mathematics Report section 3, physical PDF pages 4–6
Main empirical results Report section 4, physical PDF pages 7–10
Robustness tables and diagnostics Physical PDF pages 11–14
Complete data-preparation appendix Physical PDF pages 15–20
Complete analysis appendix Physical PDF pages 20–33
Custom clustering study and code Physical PDF pages 33–41
Portable data preparation src/01_data_processing_and_descriptives.R
MDS, PAM, robustness, and co-clustering src/02_methods_and_analysis_pipeline.R
Custom linkage implementation src/03_custom_agglomerative_hierarchical_clustering.R
Dependency installer and version reporter requirements.R
Machine-readable citation CITATION.cff

The report cover is physical PDF page 1; printed page 1 is physical page 2. Generated figures and .rds intermediates are intentionally absent because the report already packages the results and the scripts recreate them under ignored output/ and data_processed/ directories.

Reproduce the analysis

1. Clone and install dependencies

git clone https://github.com/SadraDaneshvar/call-of-the-blocs-unga-mds-clustering.git
cd call-of-the-blocs-unga-mds-clustering
make setup

make setup executes requirements.R, which installs only missing packages and prints the resolved versions. Repository verification used R 4.5.1 with unvotes 0.3.0, smacof 2.1-7, blockcluster 4.5.5, cluster 2.1.8.1, and mclust 6.1.2. Unlike a lockfile, this package list does not freeze transitive versions; the limitation is explicit below.

2. Run the complete project

make run

The target rebuilds the analytic vote table, runs MDS and PAM diagnostics, evaluates abstention sensitivity, fits the nine-model categorical co-clustering grid, and validates the custom agglomerative implementation.

For focused execution:

make data
make analysis
make custom

The latent-block grid is the computationally intensive stage. All paths resolve from the repository root, analysis programs do not install packages at runtime, and generated files remain outside version control.

3. Run dataset-free integrity checks

make validate

This parses requirements.R and every analysis program and verifies the SHA-256 digest of the canonical report.

Interpretation boundaries

  • Blocs summarize voting records; they are not alliances or normative categories. Membership captures similarity under the stated distance and period.
  • The 1990–2019 window privileges persistent structure. It can obscure enlargement, regime change, conflict, and realignment within the period.
  • Recorded absence is ambiguous. The co-clustering state Absent preserves non-observation but cannot distinguish every procedural or political reason for it.
  • Exact-match distance gives every unequal pair of states the same baseline cost. Two alternative abstention treatments demonstrate robustness, not the uniqueness of that coding.
  • Overlap weighting addresses precision, not selection. It gives heavily supported country pairs more influence but does not make participation missing at random.
  • A two-dimensional map is a compression. Stress and Shepard diagnostics quantify fit; displayed axes, orientation, and absolute coordinates have no intrinsic meaning.
  • PAM and latent-block labels are permutation-invariant. Numeric cluster IDs have meaning only inside a particular fitted result.
  • Issue labels are post-estimation annotations. They help explain resolution clusters but do not validate a unique political interpretation.
  • The tested co-clustering grid is finite. A 4-by-4 optimum means highest ICL among nine candidates, not proof that no larger block structure could fit better.
  • Package versions are recorded but not locked. Fixed seeds stabilize stochastic starts, while future dependency changes may still affect numerical results.

Data and method references

The final report contains the full academic bibliography and methodological discussion.

License

Code and repository documentation are released under the MIT License. Machine-readable citation metadata is available in CITATION.cff. The source datasets remain governed by their own upstream terms and citations.

Releases

Packages

Contributors

Languages