Unicode Frequently Asked Questions

Programming Issues

Q: How do I convert an existing application written in 'C'; so that it can handle Unicode strings?

There is no simple answer to that. The optimal solution depends on the nature of your application, the nature of the data it reads, and the nature of the APIs you are going to use. Assuming that your application currently reads and manipulates ASCII strings, the first thing to look at is the encoding form of Unicode you are going to use. [MS]

Q: When would using UTF-8 be the right approach?

If your input data and APIs use UTF-8, working directly in UTF-8 with byte strings is often a suitable choice. It avoids conversions at API boundaries and is compact for ASCII-heavy data such as HTML.

Choose an encoding based on the APIs, data, and operations your application needs. Working in UTF-8 does not require unusually large data volumes or limited text processing. Unicode libraries can provide operations such as normalization, case conversion, segmentation, and collation; use interfaces that support the encoding you choose.

Operations that interpret characters must account for multibyte sequences, while many operations on whole strings or ASCII delimiters can work directly on bytes. Measure performance and memory use with representative data if those are deciding factors.

Q: When would using UTF-16 be the right approach?

If the APIs you are using, or plan to use, are UTF-16 based, then working with UTF-16 directly is likely your best bet. Converting data for each individual call to an API is difficult and inefficient, while working around the occasional character that takes two 16-bit code units in UTF-16 is not particularly difficult (and does not have to be expensive). [MS]

Q: How about converting to UTF-32?

If your platform or i18n library supports UTF-32 (4-byte) characters, then, for similar reasons, you might want to use them instead of UTF-16. Generally, the simplest approach is to match your basic character and string datatype to that used by the APIs you need to call.

UTF-32 uses four bytes per Unicode scalar value. It can increase memory and cache requirements compared with UTF-8 or UTF-16; the effect on performance depends on the text and the operations performed.

Many libraries, such as ICU or Java, use a hybrid approach. For strings, they use UTF-16 to reduce storage, but for single-character APIs they use code points (UTF-32 values) for API simplicity. [MS]

Q: What basic datatypes do I need to use?

If you are developing for a cross-platform or cross-compiler implementation, you need to pay attention to how you define a datatype that can contain the code units of your preferred Unicode encoding form in a portable way.

On platforms with 8-bit bytes, char or unsigned char can hold UTF-8 code units. C and C++ do not require 8-bit bytes, so portable code should check CHAR_BIT and the requirements of the APIs it uses.

C11 provides char16_t and char32_t in <uchar.h>, and C++11 provides them as built-in types. Use types compatible with your APIs and check their encoding requirements. In C, the __STDC_UTF_16__ and __STDC_UTF_32__ macros indicate the corresponding Unicode encoding guarantees.

Where exact-width integer types are needed, <stdint.h> provides uint16_t and uint32_t on implementations that support those widths. Prefer standard types over compiler-specific typedefs when available.

The width and encoding of wchar_t depend on the platform. Windows uses 16-bit wchar_t for UTF-16, while other platforms commonly use a 32-bit type. Do not assume that wchar_t has a portable width or Unicode encoding.

Q: What are the porting issues I need to watch out for with UTF-8?

When porting to UTF-8, code that recognizes ASCII syntax can often be retained, because ASCII bytes represent the same characters in UTF-8 and cannot occur inside a multibyte sequence. Code that assumes one byte per character still needs review. However, watch for anything that truncates strings or buffers at places other than '\n' or '\0' or at space or syntax characters from the ASCII range. Truncations based on character counting are inherently dangerous, because UTF-8 is a multi-byte encoding. Also watch out for jumps into the middle of a string.

Many kinds of inner loop code exist for which the code does not need to be aware of the multi-byte nature of UTF-8, for example a simple copy operation like:

memcpy(d, s, len);

This copies ASCII or UTF-8 without interpreting its bytes. (len must be the byte length of the complete string including the terminating NUL.) Note that arbitrary byte-count truncation can split a UTF-8 sequence. [MS]

Q: How about issues in porting to UTF-16 or UTF-32?

If you port to UTF-16 or UTF-32 you need to make sure that you use the correct datatype (see above). If you have used char* extensively for both strings and raw data buffers, you'll have your work cut out for you in deciding which pointers need to be converted to the new data type. However, compilers can be of some help here. As you convert some of the interfaces, type mismatches should be flagged. If you are using C, try compiling with a C++ compiler, since even though you are writing C code, your type checking will generally improve.

If all the characters that your application deals with explicitly are from the BMP (and typically from just the ASCII range, U+0000..U+007F), then the semantics of your string handling may not be impacted at all. However, your code still needs to be made aware of the single/double code unit nature of UTF-16 to avoid incorrect buffer truncations or jumping into the middle of strings at incorrect locations. The same concerns apply as for multi-byte string handling, but with 16-bit code units instead of bytes.

UTF-32 represents each Unicode scalar value in one 32-bit code unit. This simplifies code point indexing, but a user-perceived character can still contain multiple code points, so grapheme-aware operations remain necessary. Memory use and processing efficiency depend on the text being processed and the operations performed. [MS]

Page edited by [AF]