Username research · A small, reproducible character check

Two usernames look the same. Are their characters actually the same?

Compare maple with mаple. Depending on your font, the difference may be almost invisible. The second example contains a Cyrillic letter where the first has a Latin a. Before connecting two profiles, check the text you actually collected.

Unicode username OSINT starts with a modest task: describing a string accurately. This guide gives you a local experiment that separates exact equality, normalization and visual resemblance. None of those results, by itself, identifies the person behind an account.

Two similar ceramic forms reveal different individual tiles under a magnifying glass
Concept illustration: a similar outline can hide a different underlying structure. The forms are not account identifiers.

Keep the source before cleaning the text

Save the public profile URL, the platform, the date observed and the copied username in your working notes. Record whether you copied an account handle, a display name or text from a screenshot. These are different fields; a decorative display name may not be usable as a login or profile address.

Keep an untouched copy. Use a second copy for comparisons. If you silently replace characters in your only record, a colleague cannot reproduce the discrepancy that prompted the check. When all you have is a screenshot, mark the transcription as uncertain: the image shows appearance, while OCR or retyping supplies a new string.

For discovery across platforms, use the broader username research workflow. The check here answers a narrower question about the characters already in front of you.

A character has a number as well as a shape

A Unicode code point is a numbered position in the character standard, usually written in a form such as U+0061. In our first pair, the ordinary Latin a is U+0061; the Cyrillic а is U+0430. Both can look familiar to an English reader. They remain different characters.

There is a second wrinkle: what looks like one letter can contain more than one code point. Our café example can store the final accented letter as one precomposed character, or as e followed by a combining accent. Looking harder at the screen will not reliably settle that distinction.

Unicode normalization provides defined ways to compare such representations. NFC handles canonical equivalence; NFKC additionally folds compatibility distinctions, including the fullwidth letters in the example below. Neither operation is a universal “find the real username” function.

Four pairs, three comparisons

We ran the following synthetic string pairs locally with Python 3.12.10 and its Unicode 15.0.0 database. No website or account was queried. “Yes” means that the two strings compare equal in that column.

Teaching pairExactly equalEqual after NFCEqual after NFKC
Latin maple / Cyrillic-a mаpleNoNoNo
caf\u00e9 / cafe\u0301NoYesYes
maple / fullwidth mapleNoNoYes
مینا / مينا, Persian Yeh versus Arabic YehNoNoNo

The useful surprise is the last row. NFKC did not turn Arabic Yeh into Persian Yeh. Normalization follows defined character rules; it does not guess which spelling a writer intended. A language-specific search expansion is a separate decision and should be labelled as one.

Run the same experiment with Python

Save the code as unicode_compare.py and run python unicode_compare.py. It uses the standard library, prints the equality results and writes the character names and runtime versions to unicode-results.json beside the script. The samples are encoded explicitly so copy-and-paste does not conceal the character choices.

"""Offline teaching fixture; no account lookup, network call or input cleanup."""
import json
import platform
import unicodedata as ud
from datetime import datetime, timezone
from pathlib import Path

fixtures = [
    ('Latin versus Cyrillic', 'maple', 'm\u0430ple'),
    ('Composed versus combining', 'caf\u00e9', 'cafe\u0301'),
    ('Width variant', 'maple', '\uff4d\uff41\uff50\uff4c\uff45'),
    ('Persian versus Arabic Yeh', '\u0645\u06cc\u0646\u0627', '\u0645\u064a\u0646\u0627'),
]

def describe(text):
    return [{'codepoint': f'U+{ord(c):04X}',
             'name': ud.name(c, 'UNNAMED')} for c in text]

rows = []
for label, left, right in fixtures:
    row = {'case': label, 'left': left, 'right': right,
           'left_characters': describe(left),
           'right_characters': describe(right),
           'equal_raw': left == right,
           'equal_NFC': ud.normalize('NFC', left) == ud.normalize('NFC', right),
           'equal_NFKC': ud.normalize('NFKC', left) == ud.normalize('NFKC', right)}
    rows.append(row)
    print(label, row['equal_raw'], row['equal_NFC'], row['equal_NFKC'])

report = {'checked_at_utc': datetime.now(timezone.utc).isoformat(),
          'python': platform.python_version(), 'unicode_database': ud.unidata_version,
          'scope': 'Four synthetic string pairs; no platform behavior or identity tested',
          'rows': rows}
Path(__file__).with_name('unicode-results.json').write_text(
    json.dumps(report, ensure_ascii=False, indent=2), encoding='utf-8')

Python’s unicodedata documentation explains the normalization and character-name functions. Check the recorded runtime version when reproducing the experiment; the version of the online documentation may differ from your installed Python.

Turn the result into a careful next step

  • Raw strings match: you have the same character sequence. Two platforms can still assign that sequence to unrelated people.
  • Raw strings differ, NFC matches: note that the stored sequences differ but are canonically equivalent. Check the platform’s actual username rules before assuming they address one account.
  • NFC differs, NFKC matches: keep both originals. A compatibility comparison helped group them for review; it did not authorize replacing a profile URL with a guessed spelling.
  • No match, but similar appearance: inspect the differing code points. Treat a possible lookalike as a research lead, not an accusation of impersonation.

Unicode’s security guidance describes confusable-character detection and its limitations. A dedicated confusable comparison is different from NFC or NFKC. Our short script does not implement that detector, test invisible-character policies or model a platform’s registration rules.

Non-Latin text is not suspicious by default. Multiple scripts and joining characters can be ordinary language use. The relevant question is whether a difference explains the particular account confusion you are investigating.

Write a finding someone else can check

For the first synthetic pair, a defensible note would read:

These two supplied strings differ at the second character: U+0061 versus U+0430. They remained unequal after NFC and NFKC in the recorded runtime. Their appearance may be similar. No conclusion about account ownership follows from this comparison.

Add the actual public source URLs and collection time in a real case. If the next question is whether one account changed its handle over time, move to username history and continuity; that needs different evidence.

When several plausible accounts remain, keep them as separate candidates in your investigation notes. OSINT Jet’s digital identity graph workflow is relevant to organizing relationships for review. A character comparison should explain why a proposed link is still uncertain, rather than quietly turning resemblance into a confirmed identity.

Published by OSINT Jet Editorial Team · 5 October 2026

Suggest a correction · نسخه فارسی