跳到正文

productdevbook

cizgile

cizgile — Zero-dependency URL slug engine. RFC 3986/3987 slugs, transliteration for 7 scripts, Unicode slugs, IRI ↔ URI, percent-encoding. Pure TypeScript, works everywhere.

README 已保存到本站,可直接阅读

Documentation snapshot

README 快照

这篇是英文原文

下面正文是项目自己的英文 README。想读全文就用浏览器自带的整页翻译: Chrome / Edge 点地址栏右侧的翻译图标,或用右键菜单里的「翻译成中文」; 手机浏览器一般在菜单里。

本页保存的是公开项目资料快照,阅读过程不需要连接 GitHub。

cizgile

Zero-dependency URL slug engine.

Turn any title into a clean URL slug — in any language — and work with URLs the way RFC 3986 and RFC 3987 describe them. Pure TypeScript, works everywhere.

Install

npm install cizgile
import { slugify } from "cizgile"

slugify("Hello, World!") // "hello-world"
slugify("İstanbul Şişli & Çığ", { locale: "tr" }) // "istanbul-sisli-ve-cig"
slugify("Straße Über Ärger", { locale: "de" }) // "strasse-ueber-aerger"
slugify("你好 World", { unicode: true }) // "你好-world"

No dependencies. ESM only. Node 20+, Bun, Deno, browsers, edge workers.

Why cizgile

  • Slugs that are correct by construction. Every ASCII slug is a valid URL path segment (RFC 3986 segment-nz-nc) — no percent-encoding needed, no . or .., no accidental scheme prefix.
  • Speaks your language. 19 locales (tr, de, da, sv, uk, bg, …) and 7 scripts (Latin, Cyrillic, Greek, Arabic, Armenian, Georgian, Dhivehi). ß → ss, İ → i, Щ → shch.
  • Unicode slugs when you want them. 你好-world stays readable, and iriToUri gives you the exact percent-encoded form for the wire.
  • A real URL toolkit underneath. Resolve, normalise, compare, validate and relativise URLs by the RFC, cross-checked against the WHATWG URL parser.
  • Small and tree-shakeable. import { slugify } ships the Latin table only; other scripts load only when you import them.

Three entry points

importwhat you get
cizgileslugify, isSlug, createSlugger, truncateSlug, decamelize, script and bidi guards
cizgile/transliteratetransliterate, per-script tables, locales, defineLocale
cizgile/uripercent-encoding, resolveUri, normalizeUri, relativize, validators, IRI ↔ URI, punycode

Slugs

Everyday use

import { slugify } from "cizgile"

slugify("Déjà Vu!") // "deja-vu"
slugify("don't stop") // "dont-stop"
slugify("v1.2.3", { preserveCharacters: ["."] }) // "v1.2.3"
slugify("Hello World", { separator: "_" }) // "hello_world"
slugify("Donald E. Knuth", { lowercase: false }) // "Donald-E-Knuth"
slugify("getHTTPResponse", { decamelize: true }) // "get-http-response"
slugify("the quick brown fox", { maxLength: 9 }) // "the-quick"
slugify("C++ & Rust", { replacements: [["C++", "cpp"]] }) // "cpp-and-rust"

Languages and scripts

Locale ids for Latin-script languages; Cyrillic locales and other scripts come from cizgile/transliterate so they only end up in your bundle when you use them.

import { slugify } from "cizgile"
import { cyrillic, greek, uk, defineLocale, de } from "cizgile/transliterate"

slugify("Çay & Simit", { locale: "tr" }) // "cay-ve-simit"
slugify("Fisch & Chips", { locale: "de" }) // "fisch-und-chips"
slugify("Ærø", { locale: "da" }) // "aeroe"
slugify("Київ", { locale: uk }) // "kyiv"
slugify("Привет мир", { transliterate: [cyrillic] }) // "privet-mir"
slugify("Καλημέρα", { transliterate: [greek] }) // "kalimera"

const swiss = defineLocale(de, { id: "de-CH", table: { ß: "ss" } })
slugify("Straße", { locale: swiss }) // "strasse"

Locale ids: az da de es fi fr hu it nb nl pt sv tr vi. Locale objects: those plus bg mk ru sr uk.

Unicode slugs

import { slugify } from "cizgile"
import { iriToUri, uriToIri } from "cizgile/uri"

const slug = slugify("Ünïcödé final ①", { unicode: true }) // "ünïcödé-final-1"
const wire = iriToUri(slug) // "%C3%BCn%C3%AFc%C3%B6d%C3%A9-final-1"
uriToIri(wire) === slug // true

Unicode slugs keep letters, digits and combining marks, are NFKC-normalised, never start with a mark and contain no invisible or bidi-control characters. Two optional guards for user-supplied titles:

slugify("pаypal", { unicode: true, scripts: "single" }) // throws — that "а" is Cyrillic
slugify("مرحبا 123", { unicode: true, bidi: "encode" }) // "%D9%85%D8%B1%D8%AD%D8%A8%D8%A7-123"

scripts applies the UTS #39 restriction levels ("single", "highly-restrictive", "moderately-restrictive", "any"); bidi enforces RFC 3987 §4.2 ("allow", "encode", "throw").

Unique slugs

import { createSlugger } from "cizgile"

const slug = createSlugger()
slug("Hello") // "hello"
slug("Hello") // "hello-2"
slug("hello-2") // "hello-2-2"  — never a duplicate
slug.reset()

Validation

import { isSlug } from "cizgile"

isSlug("hello-world") // true
isSlug("Hello World") // false
isSlug("hello_world", { separator: "_" }) // true
isSlug("你好-world", { unicode: true }) // true

isSlug accepts exactly what slugify would produce under the same options.

All options

optiondefaultwhat it does
separator"-"Joins words. Any URL-safe punctuation (- _ . ~ !$&'()*+,;= @) or "".
lowercasetruefalse keeps the original case.
unicodefalseKeep letters from every script instead of transliterating to ASCII.
locale—Language-specific rules: a locale id or a Locale object.
transliteratetruefalse skips the tables (accents still fold); an array adds script tables.
decamelizefalsefooBar → foo-bar, HTMLParser → html-parser.
replacements[][from, to] pairs applied first; spaces in to become separators.
remove/['’]/gCharacters to delete rather than turn into separators (don't → dont).
preserveCharacters[]Extra URL-safe characters to keep, e.g. ["."] for version numbers.
preserveLeadingUnderscorefalse_draft → _draft.
preserveTrailingSeparatorfalseKeep a trailing separator while the user is still typing.
maxLength—Cut at a word boundary, never inside a character (emoji sequences, combining marks). Counts UTF-16 code units like .length.
scripts"any"Unicode mode: UTS #39 mixed-script restriction level.
bidi"allow"Unicode mode: RFC 3987 §4.2 direction rule — "encode" or "throw" on violation.

The pipeline runs in this order: strip control/format characters → NFC → replacements → NFKC → decamelize → transliterate (locale → your tables → Latin → symbols → strip accents) → lowercase → remove → separators → maxLength → guards. Output is idempotent: slugify(slugify(x)) === slugify(x).

Transliteration on its own

import { transliterate, cyrillic, locales } from "cizgile/transliterate"

transliterate("Straße Ærø") // "Strasse AEro"
transliterate("Привет", { tables: [cyrillic] }) // "Privet"
transliterate("Ängsö", { locale: locales.sv }) // "Aengsoe"
transliterate("你好") // "你好" — unknown scripts are kept (use unknown: "drop" to remove)

Tables: latin symbols cyrillic cyrillicUk cyrillicBg cyrillicMk cyrillicSr greek arabic persian urdu pashto armenian georgian dhivehi, plus allScripts. Where a letter is spelled differently at the start of a word (Armenian ե, Ukrainian є ї й ю я), the capital carries the word-initial form. defineLocale and mergeTables return new objects — nothing global is ever mutated.

URL toolkit

Everything in cizgile/uri follows RFC 3986 / RFC 3987 to the letter and is tested against the RFC’s own examples and the WHATWG URL parser.

import {
  resolveUri,
  relativize,
  normalizeUri,
  equivalentUris,
  encodePathSegment,
  percentEncode,
  percentDecode,
  isUri,
  isAbsoluteUri,
  isIPv6Address,
  extractUri,
  iriToUri,
  uriToIri,
  domainToAscii,
} from "cizgile/uri"

resolveUri("http://a/b/c/d;p?q", "../../g") // "http://a/g"
relativize("http://a/b/c/d;p?q", "http://a/b/g") // "../g"
normalizeUri("HTTP://www.EXAMPLE.com:80/%7e%41/./b/../c") // "http://www.example.com/~A/c"
equivalentUris("http://example.com", "http://example.com:80/") // true
encodePathSegment("a/b?c") // "a%2Fb%3Fc"
percentEncode("À ア") // "%C3%80%20%E3%82%A2"
isAbsoluteUri("http://a/b#c") // false — fragments are not allowed in an absolute-URI
isIPv6Address("::ffff:192.0.2.1") // true
extractUri(".") // "http://a/b"
iriToUri("http://例え.jp/résumé", { host: "punycode" }) // "http://xn--r8jz45g.jp/r%C3%A9sum%C3%A9"

Full reference

Characters and percent-encoding (RFC 3986 §2) isUnreserved isReserved isGenDelim isSubDelim isPchar isSegmentNzNc isQueryChar isScheme — per code point. percentEncode(text, keep?) — UTF-8, uppercase hex. keep names a set: RFC "unreserved" "pchar" "segment-nz-nc" "path" "query" "fragment" "userinfo", WHATWG "whatwg-c0-control" "whatwg-fragment" "whatwg-query" "whatwg-special-query" "whatwg-path" "whatwg-userinfo" "whatwg-component" "form", or a predicate. percentDecode(text, { plusAsSpace }), normalizePercentEncoding(text). encodePathSegment(seg, { noColon }), encodePath(path, { relative }), encodeQuery, encodeFragment, encodeForm.

Hosts (§3.2.2) isIPv4Address isIPv6Address isIPvFuture isIPLiteral isRegName isHost parseHost parseAuthority serializeAuthority. 0x7f.0.0.1 and 2130706433 are registered names, not addresses (§7.4).

Parsing and validation (§4, Appendix A/B) parseUri serializeUri — components stay distinct from “absent”; the serializer inserts /. or ./ where the grammar requires it. isUriReference isUri isAbsoluteUri isRelativeReference classifyReference pathForm — validating parser built from the ABNF. isIriReference isIri isIunreserved isIpchar — the same for IRIs (RFC 3987 §2.2). extractUri(text) — Appendix C: strips <>, quotes, URL: prefixes, trailing punctuation and line-wrap whitespace.

Resolution (§5) resolveUri(base, ref, { strict, allowRelativeBase }) — every §5.4 example passes; strict by default (http:g stays http:g). relativize(base, target) — shortest reference that resolves back to target. removeDotSegments(path) — the literal two-buffer algorithm. mergePaths(base, refPath). isSameDocumentReference(base, ref, { normalize }).

Normalisation and comparison (§6) normalizeUri(uri, { defaultPorts, schemeBased, userinfo }) — case, percent-encoding, dot segments, default ports, empty path → /; userinfo: "strip-password" | "strip" for logs. normalizePath(path, { trailingSlash }). equivalentUris(a, b, { level, base, ignoreFragment }) — "simple", "syntax" or "scheme" (default). Never maps IRIs to URIs (RFC 3987 §5.3.1).

IRIs (RFC 3987) isUcschar isIprivate isBidiControl hasBidiControls. iriToUri(iri, { bidi, nfc, strict, host }) — percent-encodes without altering characters (§3.1 step 1c); host: "punycode" converts the domain; strict rejects characters no IRI may contain. uriToIri(uri) — decodes only what §3.2 allows, per component. punycodeEncode punycodeDecode domainToAscii domainToUnicode — RFC 3492, no dependencies.

Deliberately not implemented: RFC 6874 IPv6 zone identifiers (reverted by RFC 9844) and the network-based normalisation of §6.2.4.

For AI agents

If you are an assistant writing code with this library, these are the facts that matter:

  • Import paths: cizgile (slugs), cizgile/transliterate (tables, locales), cizgile/uri (URLs). ESM only, no default exports, no side effects, no runtime dependencies.
  • slugify(text, options?) returns "" for input with nothing usable — it never throws on ordinary text. It throws RangeError/TypeError only for invalid options (separator: "/", preserveCharacters containing the separator, a non-global remove regex, a negative maxLength) and, in unicode mode, when scripts or bidi: "throw" rejects the result.
  • The ASCII output is always a valid path segment; put it in a URL as-is. For unicode: true output, call iriToUri(slug) before putting it on the wire.
  • Cyrillic and other non-Latin scripts are opt-in: pass transliterate: [cyrillic] or a locale object such as uk from cizgile/transliterate. Without them, Cyrillic text produces "" in ASCII mode.
  • createSlugger() is the way to get unique slugs in a document or import job; do not append counters yourself.
  • Use resolveUri, normalizeUri and equivalentUris instead of string concatenation or new URL() when you need RFC behaviour (strict scheme handling, no special-scheme rewriting, no host IDNA unless you ask for it).
  • Every exported function has an explicit TypeScript signature; the .d.mts files in dist/ are the authoritative API.

Performance

Measured with bun run bench (vitest bench, Node 24, one core of a desktop CPU). Higher is better.

inputcizgile@sindresorhus/slugifyslugify (simov)
ASCII title (60 chars)396k ops/s136k147k
Latin with diacritics235k ops/s96k179k
Turkish, locale: "tr"212k ops/s99k216k
Cyrillic, transliterate: [cyrillic]179k ops/s86k188k
2.5 KB of mixed text6.4k ops/s4.3k3.4k
isSlug3.7M ops/s——

resolveUri runs at ~0.8M ops/s (the built-in URL parser: ~1M), removeDotSegments at 2M, percentEncode at 0.5M (encodeURIComponent: 3.3M — it is native), normalizeUri at 0.3M. Options objects are resolved once and cached structurally, so inline { locale: "tr" } literals cost nothing after the first call.

How it compares

inputcizgileDjango slugifyRails parameterize@sindresorhus/slugify
" Joel is a slug "joel-is-a-slugsamesamesame
"jack & jill"jack-and-jilljack-jilljack-jilljack-and-jill
"don't"dontdontdon-tdont
"fooBar"foobarfoobarfoobarfoo-bar
"snake_case"snake-casesnake_casesnake_casesnake-case
"Straße" (locale: "de")strassestraestrassestrasse
"Привет""" (opt-in tables)""""privet

decamelize is off by default (Django/Rails behaviour) and & is spelled out (sindresorhus behaviour); both are one option away.

Specifications

RFC 3986 (with errata 2033, 4547, 4789, 5428), RFC 3987, RFC 3492, RFC 8820, RFC 9844, the WHATWG URL Standard’s percent-encode sets, Unicode UTS #39 restriction levels and UAX #29 grapheme boundaries, Google Search Central’s URL guidance. The test suite runs every example those documents contain.

Development

bun install
bun run test      # oxlint, oxfmt, tsc, vitest under node, then vitest under bun
bun run build     # rolldown → dist/*.mjs + dist/*.d.mts
bun run coverage
bun run release   # bumpp: bump, tag, push — the tag publishes to npm

Credits

  • simov/slugify — the charmap + per-locale override idea and most Cyrillic, Greek, Arabic and symbol values.
  • sindresorhus/slugify and sindresorhus/transliterate — decamelize, custom replacements, the counter slugger, and the Armenian, Georgian and Dhivehi tables.
  • Django and Rails — the reference behaviours the parity tests are written against.
  • The WHATWG URL Standard — percent-encode sets and the parser every result is cross-checked with.
  • RFC 3986 by Berners-Lee, Fielding and Masinter, and RFC 3987 by Duerst and Suignard.
  • Rolldown, Oxc, Vitest, Bun and TypeScript.

License

MIT. Transliteration values are derived from simov/slugify and sindresorhus/transliterate (both MIT).

Official distribution

获取与安装

暂未发现可确认的官方软件包地址

当前 README 快照没有出现 npm、PyPI、Crates.io、pub.dev 等官方包页链接。本站不会根据仓库名称猜测下载地址。

本站不托管项目文件;需要安装时,请以项目维护者发布的官方文档为准。

使用前核验

本站保存公开资料用于阅读,不代表安全审计或功能背书。安装前请核对许可证、依赖来源和发布签名,不要直接运行来源不明的二进制文件或高权限脚本。