# Why Hangul file names fall apart into jamo, and how NFC fixes it

> A Hangul file name made on macOS can look decomposed or count as a different name elsewhere. The cause is the Unicode normalization form, checked with code.

- Canonical: https://jaemyeong.com/en/blog/macos-hangul-jaso-nfd-nfc/
- Published: 2026.06.26
- Updated: 2026.10.04
- Category: IT/기술
- Tags: #macOS, #Unicode, #Troubleshooting, #Git

When a file with a Hangul name made on macOS is sent to Windows or Linux, the jamo in the name (the letters that make up each Hangul syllable) can appear separated. It is not only a display problem. Git can detect a file whose content did not change as renamed, and in CI the resource path the code looks for may not match the actual file name. The same can happen when a file travels as a mail attachment, through a messenger, in a ZIP, or through cloud storage.

The cause is that the same character has two Unicode representations. This post confirms the difference between the two with code points, and covers how to bring file names and strings to one of them.

Errors in the code points and the examples were present in the earlier version of this post. I corrected them by running string normalization myself on October 4, 2026, on macOS 27.0.1, with Python 3.14.5, Swift 6.4, and Node.js 24.14.1. What I checked is string normalization, plus the `convmv` dry run and rename on one file in a temporary folder. The `convmv` result is in the section below. I did not run Automator or Git, and I did not compare file systems. The definition of the normalization forms comes from [UAX #15](https://www.unicode.org/reports/tr15/), and the advice on where to normalize is my own reading.

## Two representations of the same character

NFD is the canonically decomposed form of a character, and NFC is the canonically composed form. For the syllable 한 it looks like this.

- NFC: the single code point U+D55C (HANGUL SYLLABLE HAN).
- NFD: the three code points U+1112, U+1161, and U+11AB. In order, they are HANGUL CHOSEONG HIEUH, HANGUL JUNGSEONG A, and HANGUL JONGSEONG NIEUN.

This is the Python code I used to check, with its output.

```python
import unicodedata as u

print([f"U+{ord(c):04X}" for c in u.normalize("NFD", "한")])
print(u.normalize("NFC", "\u1112\u1161\u11ab"))
print(u.normalize("NFC", "\u1106\u1161\u11ab"))
print(u.normalize("NFC", "ㅎㅏㄴ"), u.normalize("NFD", "ㅎㅏㄴ"))
print([f"U+{ord(c):04X}" for c in "ㅎㅏㄴ"])
```

Running the Python code prints the following.

```text
['U+1112', 'U+1161', 'U+11AB']
한
만
ㅎㅏㄴ ㅎㅏㄴ
['U+314E', 'U+314F', 'U+3134']
```

The Hangul strings in the code are the syllable 한 and the three separately typed letters that are explained in the next section.

The earlier version gave the initial consonant as U+1106. U+1106 is HANGUL CHOSEONG MIEUM. As the third line of the output shows, composing U+1106, U+1161, and U+11AB gives 만 (U+B9CC), not 한.

## The visible separated letters are not the decomposed form

To explain the decomposed form, it is tempting to type the letters separately and show them that way. The earlier version did so. But the letters typed on a keyboard are U+314E, U+314F, and U+3134, which are compatibility jamo, a different set of characters. They are not the same characters as U+1112, U+1161, and U+11AB used in the decomposed form.

The fourth line of the output above shows the difference. Applying NFC to the three compatibility letters does not give 한, and applying NFD leaves them as they are. The file name that the earlier version used as a decomposed example, `ㄱㅓㅁㅅㅐㄱㅎㅘㅁㅕㄴ.png`, is also written in compatibility jamo, so normalization does not turn it into `검색화면.png`. The real decomposed form of `검색화면.png` contains final-consonant jamo such as U+11B7, U+11A8, and U+11AB.

For that reason, this post writes the decomposed form as code points and escapes, not as visible letters.

## Where the problem shows up

The earlier version explained that macOS stores file names in the decomposed form and that Windows and Linux use the composed form. That explanation can vary with the file system, the program that draws the screen, and the transfer tool, and I did not check it this time. The two cases below are not reproduced from real logs either. They describe paths by which the problem can happen.

With Git, the case to think about is this. A file named `검색화면.png` is added and pushed on macOS and pulled on Windows. If the normalization form changes along the way, a file with the same content can be detected as renamed. Then diffs appear with no change in content, unnecessary commits pile up, and conflicts can become more likely. The outcome can differ with the Git configuration and with whether a conversion happens in transfer.

In builds, image and Markdown paths on the web or in iOS are affected. In a Linux container such as GitHub Actions, a path string written in the composed form in code and a file name stored in the decomposed form have different bytes. So the file may not be found.

## Renaming files to NFC with convmv

Existing file names can be changed in bulk with [convmv](https://www.j3e.de/linux/convmv/). I ran the dry-run command and the apply command below on October 4, 2026 with convmv 2.06, in a temporary folder that held one `.txt` file with a decomposed name. convmv was already installed, so I did not run the install command again.

Install it with [Homebrew](https://formulae.brew.sh/formula/convmv).

```bash
brew install convmv
```

First preview what would change. This command keeps the encoding as UTF-8, converts to NFC, and searches subdirectories.

```bash
convmv -f utf8 -t utf8 --nfc -r [변경을_원하는_디렉토리_경로]
```

The bracketed Korean placeholder stands for the path of the directory to change.

Nothing is renamed at this step. The tool shows each planned rename in the form `mv "old name" "new name"` and prints `No changes to your files done. Would have converted 1 files in 0 seconds.` and `Use --notest to finally rename the files.` That differs from the wording in the earlier version of this post. Checking the file name after the run showed it was still decomposed. After checking the list, add `--notest` to apply it.

```bash
convmv -f utf8 -t utf8 --nfc -r --notest [변경을_원하는_디렉토리_경로]
```

Running it with `--notest` printed `Ready! I converted 1 files in 0 seconds.`, and I confirmed that the file name had changed to the composed form (NFC). The name went from 10 code points to 6.

### Adding it to the Finder context menu

For frequent use, a Quick Action can be made in Automator.

1. Create a new Quick Action in Automator.
2. Set the input to files and folders, and the application to Finder.
3. Add a Run Shell Script action and set Pass input to as arguments.
4. Put in the script below.
5. Save it under a name such as "Compose Hangul jamo," and it appears in the Finder context menu.

```bash
for i in "$@"; do
    /opt/homebrew/bin/convmv -f utf-8 -t utf-8 --nfc --notest "$i"
done
```

`/opt/homebrew` is the Homebrew path on Apple Silicon Macs from the M1 on. If the install location differs, the path has to be changed. This script has no `-r`, so unlike the command above it does not process the items inside a selected folder.

## Normalizing in code

If the form is unified at the boundary where file names or user input go into a database or storage, later comparisons become simple.

Swift has [precomposedStringWithCanonicalMapping](https://developer.apple.com/documentation/foundation/nsstring/precomposedstringwithcanonicalmapping) for the composed form and `decomposedStringWithCanonicalMapping` for the decomposed form. This is an example I ran, with the decomposed form written as escapes.

```swift
import Foundation

let decomposed = "\u{1112}\u{1161}\u{11AB}"   // NFD: U+1112 U+1161 U+11AB
let composed = decomposed.precomposedStringWithCanonicalMapping
let back = composed.decomposedStringWithCanonicalMapping

func scalars(_ s: String) -> String {
    s.unicodeScalars.map { String(format: "U+%04X", $0.value) }.joined(separator: " ")
}

print(scalars(decomposed))
print(scalars(composed))
print(scalars(back))
print(decomposed == composed)
print(Array(decomposed.utf8) == Array(composed.utf8))
```

The Swift code prints these code points.

```text
U+1112 U+1161 U+11AB
U+D55C
U+1112 U+1161 U+11AB
true
false
```

The fourth and fifth lines deserve attention. A Swift `String` compared with `==` treats the decomposed and composed forms as equal. Their UTF-8 bytes, however, differ. The difference stays hidden while comparing inside Swift, and appears at the moment the string is written out as bytes and used as a file name or a key.

In JavaScript, use `normalize("NFC")`.

```javascript
const decomposed = "\u1112\u1161\u11AB"; // NFD: U+1112 U+1161 U+11AB
const composed = decomposed.normalize("NFC");

const codePoints = (s) => [...s].map((c) => "U+" + c.codePointAt(0).toString(16).toUpperCase().padStart(4, "0")).join(" ");

console.log(codePoints(decomposed));
console.log(codePoints(composed));
console.log(decomposed === composed);
console.log(decomposed.normalize("NFC") === composed);
```

The Node.js code prints this.

```text
U+1112 U+1161 U+11AB
U+D55C
false
true
```

The `===` operator in JavaScript treats the decomposed and composed forms as different. They become equal only after normalization.

In Python it is `unicodedata.normalize`. This function only returns a changed string. It does not rename a file on disk.

```python
import os
import sys
import unicodedata

def sanitize_nfc(path):
    return unicodedata.normalize('NFC', path)
```

## Places to check

Using only English file names avoids the problem, but for a project that uses Hangul names, it is better to check the following.

- Whether decomposed names have entered a repository that several people use.
- Whether a build on Windows or Linux finds resources at Hangul paths.
- How Git detects renames. `core.ignorecase` is a setting about letter case, and there is no basis for saying it solves normalization.
- Whether the boundary that receives uploads normalizes and validates names.

Whether `convmv` converts every file, and whether names stay intact in transfer after the conversion, was not checked this time. The same goes for what happens when a converted name collides with an existing file, for permission problems, and for differences between file systems. What I confirmed by running is string normalization in Python, Swift, and JavaScript, plus the `convmv` dry run and rename on one file in a temporary folder.
