Summary
The client-side URL-slug preview (shown while typing a dataset/group/org title) produces an incorrect result when the title contains a character such as å, ä or ö that arrives as decomposed Unicode (a base letter followed by a separate combining mark) rather than as a single precomposed code point. This is common when the text is pasted from certain word processors, PDF extractors, or other tools that internally use decomposed Unicode.
CKAN version
2.11 (reproduced and root-caused against 2.11.6's shipped JS). Confirmed the affected file is byte-for-byte identical on current master, so this still applies there.
Steps to reproduce
- Go to "Add Dataset" (or any form with the title→URL slug live preview).
- Type a title containing å/ä/ö normally, e.g.
Kraftverk vid Åsele — note the correct preview: kraftverk-vid-asele.
- Now simulate a paste of the same visual text but with å encoded as a decomposed sequence (base "a" + U+030A COMBINING RING ABOVE) instead of the precomposed U+00E5. Easiest from the browser console:
var el = document.querySelector('#field-title');
el.value = 'Kraftverk vid åsele';
el.dispatchEvent(new Event('input', {bubbles: true}));
- Observe the live slug preview.
Expected behavior
Both inputs are the same visible text ("Kraftverk vid Åsele") and should produce the same slug: kraftverk-vid-asele.
Actual behavior
The decomposed input produces kraftverk-vid-a-sele — the combining mark is treated as an unrecognized character and turned into a stray hyphen.
Root cause
ckan/public/base/javascript/plugins/jquery.url-helpers.js's slugify() walks the input one UTF-16 code unit at a time through a fixed character map (this.map). A precomposed å (U+00E5) has an entry and maps correctly to a. A decomposed å is two separate code points: a plain ASCII a (passes through unchanged) and U+030A (no entry in the map, falls through to the default -).
Scope — which characters this affects
This is a Unicode-encoding-form bug, not a missing-character bug, so its scope is specific: it affects any character in the existing map that has a canonical Unicode decomposition into base letter + combining mark — that covers most Latin letters with diacritics (å, ä, ö, é, ü, ñ, ç, and equivalents in many other languages), since a decomposed and precomposed form of the same visible character can both be pasted in.
It does not affect letters that have no decomposed form at all, e.g. Danish/Norwegian ø/Ø, German ß, Polish ł, Icelandic þ/ð — these are atomic Unicode code points with only one possible encoding, so they can never arrive "decomposed" and were never affected by this bug in the first place (we checked: ø/Ø are already present and correctly mapped in this.map).
Note on the server-side slug
We also checked ckan.lib.munge.substitute_ascii_equivalents(), the authoritative saved-name generator server-side, expecting the same bug. It does not have this problem: it silently drops (rather than hyphenates) any character it doesn't recognize, so a decomposed combining mark simply disappears and the correct base letter survives, e.g. it already produces kraftverk-vid-asele for the same decomposed input. So the impact of this bug is limited to the live preview showing something the server would not actually save.
Suggested fix
Normalize the input string to Unicode NFC before slugifying, so precomposed and decomposed input are equivalent before the character map is applied. We have a minimal patch and a regression test ready — see the linked pull request.
Note
This issue was drafted with the assistance of an AI coding tool (Claude), under my direction — I've reviewed the repro steps and root-cause analysis myself. I'm new to contributing to CKAN, so apologies for any missteps in format or convention.
Summary
The client-side URL-slug preview (shown while typing a dataset/group/org title) produces an incorrect result when the title contains a character such as å, ä or ö that arrives as decomposed Unicode (a base letter followed by a separate combining mark) rather than as a single precomposed code point. This is common when the text is pasted from certain word processors, PDF extractors, or other tools that internally use decomposed Unicode.
CKAN version
2.11 (reproduced and root-caused against 2.11.6's shipped JS). Confirmed the affected file is byte-for-byte identical on current
master, so this still applies there.Steps to reproduce
Kraftverk vid Åsele— note the correct preview:kraftverk-vid-asele.Expected behavior
Both inputs are the same visible text ("Kraftverk vid Åsele") and should produce the same slug:
kraftverk-vid-asele.Actual behavior
The decomposed input produces
kraftverk-vid-a-sele— the combining mark is treated as an unrecognized character and turned into a stray hyphen.Root cause
ckan/public/base/javascript/plugins/jquery.url-helpers.js'sslugify()walks the input one UTF-16 code unit at a time through a fixed character map (this.map). A precomposed å (U+00E5) has an entry and maps correctly toa. A decomposed å is two separate code points: a plain ASCIIa(passes through unchanged) and U+030A (no entry in the map, falls through to the default-).Scope — which characters this affects
This is a Unicode-encoding-form bug, not a missing-character bug, so its scope is specific: it affects any character in the existing map that has a canonical Unicode decomposition into base letter + combining mark — that covers most Latin letters with diacritics (å, ä, ö, é, ü, ñ, ç, and equivalents in many other languages), since a decomposed and precomposed form of the same visible character can both be pasted in.
It does not affect letters that have no decomposed form at all, e.g. Danish/Norwegian ø/Ø, German ß, Polish ł, Icelandic þ/ð — these are atomic Unicode code points with only one possible encoding, so they can never arrive "decomposed" and were never affected by this bug in the first place (we checked: ø/Ø are already present and correctly mapped in
this.map).Note on the server-side slug
We also checked
ckan.lib.munge.substitute_ascii_equivalents(), the authoritative saved-name generator server-side, expecting the same bug. It does not have this problem: it silently drops (rather than hyphenates) any character it doesn't recognize, so a decomposed combining mark simply disappears and the correct base letter survives, e.g. it already produceskraftverk-vid-aselefor the same decomposed input. So the impact of this bug is limited to the live preview showing something the server would not actually save.Suggested fix
Normalize the input string to Unicode NFC before slugifying, so precomposed and decomposed input are equivalent before the character map is applied. We have a minimal patch and a regression test ready — see the linked pull request.
Note
This issue was drafted with the assistance of an AI coding tool (Claude), under my direction — I've reviewed the repro steps and root-cause analysis myself. I'm new to contributing to CKAN, so apologies for any missteps in format or convention.