Mapping post-Cold War UN General Assembly alignment through overlap-weighted MDS, medoid blocs, and country–resolution co-clustering
Read the final report · Build the vote matrix · Run MDS, PAM, and co-clustering · Inspect the custom linkage code
Original academic project title: Call of the Blocs
A UN General Assembly roll call is both a diplomatic event and an entry in a very large relational table. Viewed across three decades, that table can reveal durable country-to-country alignment; viewed resolution by resolution, it can reveal coalitions that reorganize with the issue. Call of the Blocs models both structures without pretending that an unrecorded vote is an observed disagreement.
The first lens compares countries only on roll calls where both have a recorded vote, weights each comparison by how much shared evidence supports it, embeds the resulting dissimilarities with ratio-SMACOF multidimensional scaling, and discovers broad blocs with PAM. The second lens returns to the full country-by-resolution table and jointly clusters both sides with a categorical latent block model, preserving absence as its own state.
| 176 countries after participation filters |
2,609 post-Cold War roll calls |
420,230 recorded Yes, Abstain, and No votes |
15,400 country pairs with nonzero overlap |
This repository is a compact technical guide to the analysis. The report remains the canonical evidence package with every figure, table, appendix, and academic reference; the public R programs make the full computational path inspectable and runnable.
- The substantive question
- One voting relation, two analytical lenses
- Data lineage and analytic sample
- Shared evidence before distance
- Weighted ratio-SMACOF MDS
- Country blocs: PAM on the observed dissimilarities
- Resolution-sensitive structure: categorical co-clustering
- Custom agglomerative clustering: a separate algorithm check
- Results ledger
- Artifact map
- Reproduce the analysis
- Interpretation boundaries
- Data and method references
- License
How is post-Cold War UNGA voting similarity organized across countries, and what recurring coalitions become visible only when the resolution dimension is retained?
The question has a built-in missing-data problem. A country that has no recorded vote on a resolution supplies neither Yes, Abstain, nor No. Pairwise agreement must therefore be calculated over the intersection of two countries' observed roll calls, and comparisons supported by more shared roll calls should carry more weight.
The project is descriptive and unsupervised. It identifies structure in recorded votes; it does not predict future positions, infer governments' motivations, rank countries normatively, or estimate a causal effect of bloc membership.
The primitive object is bipartite: delegations connect to resolutions through categorical vote states. A country projection and a block model answer different questions about the same relation.
graph LR
subgraph Delegations
I((Country i))
J((Country j))
end
subgraph Roll_calls
R1{{Resolution 1}}
R2{{Resolution 2}}
R3{{Resolution 3}}
end
I -- Yes --> R1
J -- Yes --> R1
I -- No --> R2
J -- Abstain --> R2
I -- Yes --> R3
R1 --> S[Shared evidence: agreement]
R2 --> T[Shared evidence: disagreement]
R3 -.-> U[Excluded: no recorded vote for country j]
S --> P[Country projection: dissimilarity and overlap]
T --> P
I --> B[Full bipartite table]
J --> B
R1 --> B
R2 --> B
R3 --> B
B --> L[Country and resolution latent blocks]
classDef country fill:#E7F5FA,stroke:#0077A3,color:#12313D,stroke-width:2px
classDef vote fill:#FFF1D6,stroke:#B66A00,color:#3F2B0A,stroke-width:2px
class I,J country
class R1,R2,R3 vote
| Global country lens | Resolution-sensitive lens |
|---|---|
| Collapse shared voting records into one country-by-country dissimilarity matrix. | Keep the complete country-by-resolution categorical table. |
| Weight each pair by its number of jointly observed roll calls. | Encode non-recorded participation as a distinct Absent state. |
| Use MDS for continuous geometry and PAM for broad blocs. | Estimate country groups, resolution groups, and block-specific vote probabilities jointly. |
| Answer: who tends to vote alike across the period? | Answer: which coalitions recur for which sets of resolutions? |
This split prevents an attractive global map from becoming the only model of alignment.
The scripts load un_votes, un_roll_calls, and un_roll_call_issues from the unvotes R package. That package distributes a cleaned form of the Voeten, Strezhnev, and Bailey UN voting archive, whose research-data record is available through Harvard Dataverse. The UN Digital Library voting-data record provides a separate official reference point for General Assembly voting records; it is not silently substituted for the versioned package input.
The construction is deterministic:
- keep recorded
Yes,Abstain, andNooutcomes; - restrict roll calls to 1990–2019;
- retain countries with at least 1,500 recorded votes in that window;
- retain roll calls with at least 120 participating countries;
- preserve issue labels for post-estimation interpretation, never as clustering inputs.
| Stage | Countries | Roll calls | Recorded votes |
|---|---|---|---|
| Post-Cold War window before stability filters | 194 | 2,650 | 444,195 |
| Final analytic sample | 176 | 2,609 | 420,230 |
| Final recorded state | Count | Share of recorded votes |
|---|---|---|
| Yes | 341,442 | 81.25% |
| Abstain | 51,012 | 12.14% |
| No | 27,776 | 6.61% |
The preparation stage writes intermediate .rds objects under ignored data_processed/; no copied source-data binary is committed. See report section 2, physical PDF pages 2–3.
Let i and j index countries and r index roll calls. For the country-level projection, recorded votes are coded as
with V_ir missing when country i has no recorded vote on resolution r. Define the observation indicator
The number of jointly observed roll calls is
and exact agreements over that shared evidence are
The baseline dissimilarity and its reliability weight are
Thus δ_ij = 0 means complete exact agreement on the countries' shared recorded votes, while δ_ij = 1 means complete disagreement on that shared set. No country pair has zero overlap in the final sample.
| Overlap diagnostic | Roll calls |
|---|---|
| Minimum | 842 |
| 10th percentile | 1,727 |
| Median | 2,277 |
| 90th percentile | 2,541 |
| Maximum | 2,609 |
The overlap weight does not alter δ_ij; it changes how strongly that value influences the embedding. A distance based on 2,500 common votes contributes more evidence than one based on 850.
Multidimensional scaling assigns country i a coordinate x_i in p dimensions and compares Euclidean map distances with observed voting dissimilarities:
Ratio MDS estimates one common scale factor through the origin:
The weighted raw-stress loss is
and the reported scale-free Stress-1 diagnostic is
SMACOF minimizes stress through iterative majorization. The implementation fits dimensions 1–4 with six random starts each, retains two dimensions at the empirical elbow, and refits the two-dimensional ratio model from 20 deterministic starts. Each fit allows at most 1,000 majorization iterations with tolerance 1e-6; all random starts derive from seed 1363.
| Two-dimensional fit diagnostic | Exact result |
|---|---|
| Stress-1 | 0.103895 |
| Fitted ratio slope | 0.067790 |
| Squared weighted Shepard correlation | 0.975341 |
The last quantity is the squared weighted correlation between fitted disparities and configuration distances. It is a descriptive Shepard-fit measure, not an out-of-sample predictive coefficient of determination.
MDS coordinates are invariant to translation, rotation, and reflection. Neither axis direction nor sign has intrinsic political meaning; only inter-country distances, neighborhoods, and broad regions are interpreted. See report section 3 and Figure 2, physical PDF pages 4–7.
PAM, or k-medoids, operates directly on the 176-by-176 dissimilarity matrix. For candidate medoids m_1 through m_K, it minimizes
Unlike a centroid, each medoid is an observed country. The analysis fits K = 2, ..., 10 and chooses the largest average silhouette width.
| PAM result | Exact value |
|---|---|
| Selected number of blocs | 2 |
| Average silhouette width | 0.650491 |
| Bloc sizes | 124 and 52 countries |
| Medoids | Burkina Faso and Slovakia |
Cluster numbers are arbitrary identifiers. The medoids are representative observed voting profiles under the fitted dissimilarity—not political leaders, normative centers, or permanent bloc labels.
| Dissimilarity treatment | Selected K |
Best average silhouette | ARI against baseline |
|---|---|---|---|
| Exact-match baseline | 2 | 0.650 | 1.000 |
| Soft abstention costs | 2 | 0.706 | 1.000 |
| Recode Abstain as No | 2 | 0.689 | 0.977 |
All variants preserve the two-bloc solution. The adjusted Rand index shows that the membership partition is identical under soft costs and changes only slightly under the Abstain-as-No coding.
The ten maximally polarizing labeled roll calls in the report show a 100-percentage-point gap in Yes shares between blocs. Human-rights resolutions dominate this particular extreme set, but issue metadata is attached only after clustering. See report section 4 and Appendix Tables 2 and 4, physical PDF pages 7–12.
A single country distance averages across all resolutions. The second lens restores the complete bipartite table and represents non-recorded participation explicitly:
Country i belongs to row cluster z_i; resolution r belongs to column cluster u_r. Inside country cluster k and resolution cluster g, the categorical block probabilities are
The model grid evaluates all nine combinations with two, three, or four country clusters and two, three, or four resolution clusters. Every fit receives a deterministic model-specific seed. Integrated Completed Likelihood selects the most supported partition in the tested grid.
| Co-clustering result | Exact value |
|---|---|
| Selected country clusters | 4 |
| Selected resolution clusters | 4 |
| Selected ICL | -260,386.5 |
| Human-rights share in labeled Resolution Cluster 2 | 86.8% |
The fitted blocks show that Yes, Abstain, No, and Absent probabilities vary across resolution groups. Other discovered column clusters combine human rights, arms control and disarmament, the Palestinian conflict, colonialism, and development in different proportions. These labels interpret the model after estimation; they do not supervise it.
The result is deliberately richer than the global PAM partition: strong broad alignment coexists with issue-dependent coalition and participation patterns. See report sections 3–4 and Figure 5, physical PDF pages 6–10.
The repository also implements hierarchical agglomerative clustering from first principles on a transparent ten-point dissimilarity matrix. At each step it merges the closest active clusters and updates their distance using one of three rules:
The code maintains the active distance matrix, breaks exact ties deterministically, records the full partition at every merge, and validates every possible dendrogram cut against base R's hclust().
| Benchmark diagnostic | Single | Complete | Average |
|---|---|---|---|
| Maximum absolute merge-height gap | 0 | 0 | 0 |
| Minimum ARI across all tested cuts | 1.000 | 1.000 | 1.000 |
| Pairwise partition agreement | 100% | 100% | 100% |
| Cophenetic correlation | 0.9954866 | 0.9955481 | 0.9956034 |
This is an algorithm-validation exercise, separate from the substantive UNGA pipeline. The ten-point matrix makes each merge and encoder row inspectable. The full exercise begins at physical PDF page 33.
| Layer | Estimand or decision | Result | What it establishes |
|---|---|---|---|
| Pairwise evidence | Country overlap | 842–2,609 roll calls | Every distance is based on shared recorded votes; support varies materially. |
| Continuous geometry | Two-dimensional ratio-SMACOF | Stress-1 0.103895 | A compact map retains the major pairwise structure. |
| Broad grouping | Silhouette-selected PAM | K = 2, silhouette 0.650491 |
A strong global two-region partition is present. |
| Robustness | Abstention alternatives | ARI 0.977–1.000 | The broad split is not an artifact of one abstention coding. |
| Relational grouping | Categorical latent blocks | 4 × 4, ICL -260,386.5 |
Coalition and absence profiles vary by resolution group. |
| Algorithm verification | Manual linkage versus hclust() |
Exact partitions and merge heights | The custom clustering implementation reproduces the benchmark. |
The empirical conclusion is layered: a strong global two-bloc geometry exists, but it does not exhaust the resolution-specific structure of UNGA voting.
| Question or artifact | Canonical location |
|---|---|
| Research question and design | Report section 1, physical PDF page 2 |
| Data source, coding, and filters | Report section 2, physical PDF pages 2–3 |
| Dissimilarity, MDS, PAM, and latent-block mathematics | Report section 3, physical PDF pages 4–6 |
| Main empirical results | Report section 4, physical PDF pages 7–10 |
| Robustness tables and diagnostics | Physical PDF pages 11–14 |
| Complete data-preparation appendix | Physical PDF pages 15–20 |
| Complete analysis appendix | Physical PDF pages 20–33 |
| Custom clustering study and code | Physical PDF pages 33–41 |
| Portable data preparation | src/01_data_processing_and_descriptives.R |
| MDS, PAM, robustness, and co-clustering | src/02_methods_and_analysis_pipeline.R |
| Custom linkage implementation | src/03_custom_agglomerative_hierarchical_clustering.R |
| Dependency installer and version reporter | requirements.R |
| Machine-readable citation | CITATION.cff |
The report cover is physical PDF page 1; printed page 1 is physical page 2. Generated figures and .rds intermediates are intentionally absent because the report already packages the results and the scripts recreate them under ignored output/ and data_processed/ directories.
git clone https://github.com/SadraDaneshvar/call-of-the-blocs-unga-mds-clustering.git
cd call-of-the-blocs-unga-mds-clustering
make setupmake setup executes requirements.R, which installs only missing packages and prints the resolved versions. Repository verification used R 4.5.1 with unvotes 0.3.0, smacof 2.1-7, blockcluster 4.5.5, cluster 2.1.8.1, and mclust 6.1.2. Unlike a lockfile, this package list does not freeze transitive versions; the limitation is explicit below.
make runThe target rebuilds the analytic vote table, runs MDS and PAM diagnostics, evaluates abstention sensitivity, fits the nine-model categorical co-clustering grid, and validates the custom agglomerative implementation.
For focused execution:
make data
make analysis
make customThe latent-block grid is the computationally intensive stage. All paths resolve from the repository root, analysis programs do not install packages at runtime, and generated files remain outside version control.
make validateThis parses requirements.R and every analysis program and verifies the SHA-256 digest of the canonical report.
- Blocs summarize voting records; they are not alliances or normative categories. Membership captures similarity under the stated distance and period.
- The 1990–2019 window privileges persistent structure. It can obscure enlargement, regime change, conflict, and realignment within the period.
- Recorded absence is ambiguous. The co-clustering state
Absentpreserves non-observation but cannot distinguish every procedural or political reason for it. - Exact-match distance gives every unequal pair of states the same baseline cost. Two alternative abstention treatments demonstrate robustness, not the uniqueness of that coding.
- Overlap weighting addresses precision, not selection. It gives heavily supported country pairs more influence but does not make participation missing at random.
- A two-dimensional map is a compression. Stress and Shepard diagnostics quantify fit; displayed axes, orientation, and absolute coordinates have no intrinsic meaning.
- PAM and latent-block labels are permutation-invariant. Numeric cluster IDs have meaning only inside a particular fitted result.
- Issue labels are post-estimation annotations. They help explain resolution clusters but do not validate a unique political interpretation.
- The tested co-clustering grid is finite. A 4-by-4 optimum means highest ICL among nine candidates, not proof that no larger block structure could fit better.
- Package versions are recorded but not locked. Fixed seeds stabilize stochastic starts, while future dependency changes may still affect numerical results.
- CRAN,
unvotes: United Nations General Assembly Voting Data, the authoritative package record for the versioned data objects loaded by the scripts. - E. Voeten, A. Strezhnev, and M. Bailey, United Nations General Assembly Voting Data, Harvard Dataverse.
- United Nations Digital Library, UN General Assembly voting-data record, an official reference for voting records and documentation.
- M. A. Bailey, A. Strezhnev, and E. Voeten, “Estimating Dynamic State Preferences from United Nations Voting Data”, Journal of Conflict Resolution, 2017.
- J. de Leeuw and P. Mair, “Multidimensional Scaling Using Majorization: SMACOF in R”, Journal of Statistical Software, 2009.
- P. J. Rousseeuw, “Silhouettes: A graphical aid to the interpretation and validation of cluster analysis”, Journal of Computational and Applied Mathematics, 1987.
- G. Govaert and M. Nadif, “An EM Algorithm for the Block Mixture Model”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 2005.
- L. Hubert and P. Arabie, “Comparing partitions”, Journal of Classification, 1985.
The final report contains the full academic bibliography and methodological discussion.
Code and repository documentation are released under the MIT License. Machine-readable citation metadata is available in CITATION.cff. The source datasets remain governed by their own upstream terms and citations.