Skip to content

🎨 Changed and improved the slug generation and validation - #860

Open
dittnamn wants to merge 4 commits into
TryGhost:mainfrom
dittnamn:slug-upgrade
Open

🎨 Changed and improved the slug generation and validation#860
dittnamn wants to merge 4 commits into
TryGhost:mainfrom
dittnamn:slug-upgrade

Conversation

@dittnamn

Copy link
Copy Markdown

ref towards TryGhost/Ghost#3224 etc.
In line with the changes in TryGhost/SDK#1011 a change in the validator is necessary to allow unicode slugs to be used. The validator now allows all Unicode letters and numbers, along with spaces, underscores and dashes that can be used as slug separators.

Along with the changes in the slugify() function, equivalent changes were made to the security.safe() function to allow it to use the same parameters as slugify() and to make sure the tests will still work with the new transliteration library.

However, the security.safe() function is barely used inside Ghost, and could easily be replaced to use slugify() directly. A TODO notice has been added, so that it could potentially be removed in the future.

ref towards TryGhost/Ghost#3224 etc.
In line with the changes in TryGhost/SDK#1011 a
change in the validator is necessary to allow unicode slugs to be used.
The validator now allows all Unicode letters and numbers, along with
spaces, underscores and dashes that can be used as slug separators.

Along with the changes in the slugify() function, equivalent changes
were made to the security.safe() function to allow it to use the same
parameters as slugify() and to make sure the tests will still work with
the new transliteration library.

However, the security.safe() function is barely used inside Ghost, and
could easily be replaced to use slugify() directly. A TODO notice has
been added, so that it could potentially be removed in the future.
@coderabbitai

coderabbitai Bot commented Jul 30, 2026

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 714fd6e7-56f9-4937-ba9e-3686254b8a80

📥 Commits

Reviewing files that changed from the base of the PR and between ac2ad0b and f0582c0.

📒 Files selected for processing (2)
  • packages/validator/lib/validator.js
  • packages/validator/test/internals.test.js
🚧 Files skipped from review as they are similar to previous changes (2)
  • packages/validator/lib/validator.js
  • packages/validator/test/internals.test.js

Walkthrough

The security safe helper now forwards unicodeSlugs and slugSeparator to slugify, alongside requiredChangesOnly, with updated slugging tests. The validator’s isSlug check now accepts Unicode letters, numbers, spaces, underscores, hyphens, and limited combining marks while rejecting leading or excessive combining marks and emoji. Validator tests cover these cases.

Estimated code review effort: 3 (Moderate) | ~20 minutes

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly describes the main changes to slug generation and validation.
Description check ✅ Passed The description directly explains Unicode slug support, validator changes, and updates to slugify and security.safe.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@packages/security/lib/string.js`:
- Around line 14-17: Align the options passed by the slugify call in the string
helper with the resolved `@tryghost/string`@0.3.5 API: either update or patch that
dependency to support unicodeSlugs and slugSeparator, or remove those
unsupported options and retain only the options it honors, including
requiredChangesOnly.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: d1ef6030-f681-49df-b4e0-6fc0e0be53cd

📥 Commits

Reviewing files that changed from the base of the PR and between 1600f6d and 658f179.

📒 Files selected for processing (4)
  • packages/security/lib/string.js
  • packages/security/test/string.test.js
  • packages/validator/lib/validator.js
  • packages/validator/test/internals.test.js

Comment on lines +14 to +17
return slugify(string, {
requiredChangesOnly: options.importing === true,
unicodeSlugs: options.unicodeSlugs === true,
slugSeparator: options.slugSeparator

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
node <<'NODE'
const fs = require('fs');
const path = require('path');

const entry = require.resolve('`@tryghost/string`');
const packageJson = JSON.parse(
    fs.readFileSync(path.join(path.dirname(entry), '..', 'package.json'), 'utf8')
);
const {slugify} = require('`@tryghost/string`');

const unicode = slugify('Ett smörgåsbord från Sydkorea: 스뫼르고스보르드', {
    unicodeSlugs: true
});
const separated = slugify('Herr./Klaus-Jürgen_44', {
    slugSeparator: ' '
});

console.log({version: packageJson.version, unicode, separated});

if (unicode !== 'ett-smörgåsbord-från-sydkorea-스뫼르고스보르드') process.exit(1);
if (separated !== 'herr klaus-jurgen_44') process.exit(1);
NODE

Repository: TryGhost/framework

Length of output: 957


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== dependency declarations =="
git ls-files | rg '(^|/)(package\.json|pnpm-lock\.yaml|package-lock\.json|yarn\.lock)$' | while read -r f; do
  echo "--- $f"
  rg -n '"`@tryghost/string`"|`@tryghost/string`|slugSeparator|unicodeSlugs' "$f" -C 2 || true
done

echo
echo "== relevant source =="
fd -a 'package.json$' . | while read -r f; do
  case "$f" in
    */node_modules/*) continue ;;
  esac
  if rg -q '`@tryghost/string`' "$f"; then
    echo "--- $f"
    sed -n '/@tryghost\/string/,+2p' "$f"
  fi
done

echo
echo "== read package file =="
if [ -f packages/security/lib/string.js ]; then
  cat -n packages/security/lib/string.js | sed -n '1,80p'
fi

echo
echo "== lockfile snippets =="
if [ -f pnpm-lock.yaml ]; then
  rg -n 'tryghost/string|/(`@tryghost/string`):|slugSeparator|unicodeSlugs' pnpm-lock.yaml -C 3 || true
fi

Repository: TryGhost/framework

Length of output: 4673


🌐 Web query:

@tryghost/string 0.3.5 npm source slugify unicodeSlugs slugSeparator

💡 Result:

The package @tryghost/string version 0.3.5 is a utility library maintained by the Ghost Foundation, primarily used within the Ghost publishing platform [1][2][3]. Regarding your specific search terms: 1. slugify: The package provides slugification functionality used to convert strings into URL-friendly slugs [4][5]. In modern versions of Ghost, the slugification logic is designed to handle character encoding specifically to ensure compatibility across browsers [4][6]. 2. unicodeSlugs: There is no native option or parameter named unicodeSlugs in the @tryghost/string library. Ghost's default slugification process often strips non-ASCII characters or uses encoding methods to produce predictable, ASCII-compatible slugs [4][7][6]. Users seeking to support non-Latin (Unicode) characters in slugs often find it necessary to manually modify the underlying slugify logic in their Ghost installation, as there is no built-in configuration option for this [5][7][6]. 3. slugSeparator: Similar to unicodeSlugs, there is no explicit parameter named slugSeparator exposed in the public interface of this package to configure the separator character (which defaults to a hyphen) [4]. Because @tryghost/string is a private or internal-facing utility for the Ghost project, it lacks public, comprehensive API documentation [8][9]. The implementation of slugification is typically found in the core codebase of Ghost, where it is used to process titles and other strings into slugs [4][5][7]. If you require specific customization, you would need to inspect the source code within the Ghost repository or the specific version of the package installed in your node_modules directory [4][7].

Citations:


🌐 Web query:

npmx.dev @tryghost/string v0.3.5 slugify.js unicodeSlugs slugSeparator

💡 Result:

The @tryghost/string package, including version 0.3.5 [1][2], is a utility library maintained by the Ghost Foundation, commonly used within the Ghost ecosystem to handle string operations such as URL slug generation [3][4][5]. Regarding the specific terms mentioned: slugify.js: This is a core component within the @tryghost/string package (and historically within other internal Ghost packages like @tryghost/kg-utils) that provides the logic for converting strings into URL-safe slugs [6][3][7]. The implementation typically involves replacing reserved URL characters, handling whitespace, and normalizing the string to lowercase [6]. unicodeSlugs and slugSeparator: These terms are not explicit configuration options or variables found within the standard slugify.js implementation of @tryghost/string. The library’s slugification logic is generally prescriptive rather than highly configurable via these specific names [6]. While different versions of Ghost have evolved their slugification strategies—moving from aggressive character replacement (using dashes as separators) to URL encoding in newer versions—they do not expose a public API using the specific identifiers "unicodeSlugs" or "slugSeparator" [8][9]. If you are looking for advanced configuration for slug generation (such as custom separators or specific Unicode handling), libraries like @sindresorhus/slugify are often used in the broader JavaScript ecosystem as they provide explicit options for features like separators and Unicode transliteration [10].

Citations:


🌐 Web query:

"app.unpkg.com" "@tryghost/string@0.3.5" "unicodeSlugs" "slugSeparator"

💡 Result:

The package @tryghost/string version 0.3.5 is a utility library maintained by the Ghost project, primarily used for string processing tasks such as generating URL slugs [1]. Technical analysis of the library and its usage within the Ghost ecosystem indicates the following regarding the terms mentioned: 1. unicodeSlugs: There is no native configuration option named unicodeSlugs in the @tryghost/string library [1][2]. Ghost's default slugification behavior historically strips non-ASCII characters or encodes them [3][2]. For example, in newer versions (4.x+), the slugify function uses encodeURIComponent to handle non-Latin characters, which results in percent-encoded strings rather than retaining the original Unicode characters in the URL slug [3][2]. Users requiring full Unicode support in slugs often have to implement custom modifications to the library's slugify logic or validator expressions to override these constraints [4]. 2. slugSeparator: Similarly, the library does not expose a public configuration option named slugSeparator to change the default character used to separate words in slugs [1]. The internal implementation of the slugify function defaults to using a hyphen (-) as the separator [3][2]. If you are encountering these terms, they are likely references to custom implementations, intended features in a different context, or modifications made by users attempting to bypass the library's default ASCII-focused or hyphen-delimited behavior [4]. For standard usage, the library follows the patterns defined in the Ghost codebase, which utilize a hardcoded hyphen separator and sanitize inputs to maintain URL compatibility [3][2].

Citations:


🌐 Web query:

@tryghost/string package 0.3.5 npm source "slugSeparator" "unicodeSlugs"

💡 Result:

The @tryghost/string package, specifically version 0.3.5, is an internal utility package maintained by the Ghost organization for string manipulation, including slug generation [1][2]. Regarding your specific queries: 1. slugSeparator: The @tryghost/string package uses a hardcoded hyphen (-) as the slug separator in modern versions (Ghost 4.x and later) [3][4]. The logic within the package is designed to replace whitespace with hyphens and strip or URL-encode other characters [3][4]. There is no configurable slugSeparator option exposed in the public API of this package [3]. 2. unicodeSlugs: The package historically prioritizes ASCII-friendly URLs [5][6]. While Ghost's slugification logic has evolved, it typically performs operations like removing non-word characters or using encodeURIComponent to handle non-ASCII input [3][4]. It does not natively support "unicode slugs" in the sense of keeping raw UTF-8 characters in the URL path; instead, it tends to strip non-ASCII characters or encode them [3][4]. Users seeking to support non-Latin languages in URLs have historically had to patch the package or modify core Ghost files to bypass these restrictions [6]. Because this is an internal package used by the Ghost platform, documentation is minimal and tailored to the platform's internal requirements [2][7]. It is not intended for general public use as a standalone library [7].

Citations:


Make the supported slugify options match the workspace dependency.

packages/security/lib/string.js passes unicodeSlugs and slugSeparator, but the lockfile resolves @tryghost/string to 0.3.5, and that build exposes unidecode backing rather than those custom options. Either upgrade/patch the workspace dependency to support these options or keep passing only the options that the resolved @tryghost/string@0,3.5 implementation honors.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/security/lib/string.js` around lines 14 - 17, Align the options
passed by the slugify call in the string helper with the resolved
`@tryghost/string`@0.3.5 API: either update or patch that dependency to support
unicodeSlugs and slugSeparator, or remove those unsupported options and retain
only the options it honors, including requiredChangesOnly.

dittnamn added a commit to dittnamn/Ghost that referenced this pull request Jul 30, 2026
fixes TryGhost#3224
The slugs generated by Ghost have had a lot of issues in how they've
been transliterated. Through a change in @tryghost/string a new
transliteration library, any-ascii, has been proposed to replace
unidecode, see TryGhost/SDK#1011

However, to get a fully internationalized site along with SEO optimized
URL:s, that change isn't fully enough. The current change adds an option
in Labs where it's possible to enable Unicode slugs instead of
performing the transliteration. The Unicode characters in use are
limited to those that are registered as letters and numbers, which means
that emojis, special characters, etc. still will be removed.

The only "major" change that has been done in Ghost to make this work is
actually just to normalize the URL slugs before looking them up in the
database. The rest of the code changes are mostly settings to turn the
new feature on or off, and to pass the setting along all the way to the
slugify() function. Due to the rest of the system already being in
Unicode, everything seems to work as it's supposed to.

Instead of using safe() from @tryghost/security, the slug generation has
also been changed to utilize slugify() from @tryghost/string directly
instead. The routing controllers for rss, collection and channel were
previously using safe() as well to slugify the lookups before going to
the router, but the entry controller didn't do this. As the router
already checks the slugs with isSlug() before passing them further, this
was unnecessary and had the potential to break lookups, so it was simply
removed.

In addition to using Unicode slugs, there's also an option added to
switch which separator the slugs use. Previously, dashes (-) were used
as the hardcoded default, but for URL readability, underscores (_) could
be preferred, like in the URL:s Wikipedia use. There's also an option
for using spaces ( ), but this might still be seen as a bit foreign in
URL:s. Out of the major browsers, it's currently just Firefox that show
these as spaces by default, with Safari doing it in some situations.
Other browsers can show the spaces as %20.

When activated, the Unicode slugs aren't added to member tags,
newsletter, services, integrations or benefits, as these aren't user
facing.

Note that for this change to work, the changes in
TryGhost/SDK#1011 and
TryGhost/framework#860 are required.
Combining marks are now permitted to be used in a valid slug, but to
avoid misuse, marks in the beginning of a slug or more than three to a
letter in other places means the slug is invalid.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@packages/validator/lib/validator.js`:
- Around line 50-54: The slug validation regex in the validator must reject
combining marks that follow separators or any non-letter/non-number character,
not only marks at the beginning. Update the pattern used by the slug validation
return so marks are permitted only after valid letters or numbers, while
preserving the existing maximum of three combining marks per letter and allowed
slug characters.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 33a10cf8-b2e3-47bb-85b6-bed259850f7b

📥 Commits

Reviewing files that changed from the base of the PR and between 658f179 and ac2ad0b.

📒 Files selected for processing (2)
  • packages/validator/lib/validator.js
  • packages/validator/test/internals.test.js
🚧 Files skipped from review as they are similar to previous changes (1)
  • packages/validator/test/internals.test.js

Comment thread packages/validator/lib/validator.js Outdated
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant