Growth by Design Blog | Marketing Strategy & Business Growth

The Formatting Penalty: Two Habits That Make ChatGPT Skip You

Written by Izzy Gregorio | Sep 9, 2026, 2:00:00 PM

Data Report · AI Citations

The Formatting Penalty: Two Habits That Make ChatGPT Skip You

Two of the most common formatting habits on LinkedIn make your content measurably less likely to be cited. One is completely invisible to the person using it. You will not see a drop in likes. Your post will look better than the plain-text version beside it, and it will be cited 58 percent less often.

By Izzy Gregorio  ·  Updated August 2026  ·  12 min read

 

In short

Unicode bold text costs 58 percent of your ChatGPT citation odds while earning 12 percent more reactions. Link-in-comments costs 31 percent, but buys a 144 percent lift for the destination, which can be a deliberate trade. Reaction count predicts almost nothing. Technical specificity is the dominant positive signal at 77 percent. That figure is not ours; it comes from a Scrunch study published in June 2026.

The premise

 

You are playing two games and measuring one

Reach content and citation content are graded by different judges, and the judges disagree.

The LinkedIn feed rewards emotion, rhythm, and pattern recognition. Short lines. A hook that stops the thumb. A story with a turn in it. The system underneath is optimizing for dwell time and reaction volume, and it has trained an entire profession to write a specific way.

A generative engine is not scrolling. It is answering a question somebody typed. It retrieves candidate sources, decides which actually contain the answer, and cites the ones it trusts. It has no concept of a good hook. It cannot see your reaction count in any way that matters.

So most teams run a reach playbook and quietly hope citations come along for the ride. Sometimes they do. On two specific tactics they actively do not, and the trade runs the wrong way unless you chose it on purpose.

This matters more than it used to. Scrunch estimates AI platforms cite LinkedIn content roughly 8 million times per week in the US for industry and commercial prompts, growing 13 percent month over month as of Q1 2026.

OtterlyAI, analyzing 1,310,455 LinkedIn AI citations across six platforms, found LinkedIn's share of all social citations in AI search rose from 7.8 percent in January 2026 to 11.7 percent in May. Close to one in eight. If your buyers ask an assistant about your category, LinkedIn content is shaping part of that answer. The only question is whose.

 

Provenance

 

Where these numbers come from

Conspicuouz Creative Group (CZ Creative Group) did not run this study. Scrunch did, and we want that stated plainly before you repeat any of these figures to a board.

Here is the methodology, so you can weigh it yourself.

Study Scrunch, "LinkedIn posts that robots can't resist," published June 5, 2026, by Michael Iannelli, Principal Data Scientist
Observation set 12,000 LinkedIn post observations from ChatGPT, January 15 to April 15, 2026
Analysis set Approximately 4,000 posts scored for citations and reactions
Attributes tested 21 content dimensions, annotated by a language model at high reasoning effort
Method Double machine learning, estimating each dimension while controlling for the other 20
Controls Author audience size, post age, prompt difficulty

One design choice deserves credit. The dataset includes posts ChatGPT considered relevant and then declined to cite. That is the comparison you want. Without it you are only looking at winners and guessing at why.

Three limits worth holding onto. The analysis is ChatGPT-only by design, because that is where the sample was strongest. The dimension scoring was done by a language model rather than human coders. And a 90-day window in a category that changes monthly is a snapshot, not a law.

Finding one

The unicode penalty

LinkedIn's post composer has no bold button. So the workaround spread: paste your text into a formatter, and it swaps each letter for a lookalike character from the Unicode Mathematical Alphanumeric Symbols block, the range U+1D400 to U+1D7FF.

Those characters were added to Unicode so mathematicians could write formulas where a bold variable means something different from an italic one. They are not letters wearing a bold coat. They are separate characters that happen to look like letters.

To your eye, the word reads normally. To a retrieval system that does not normalize them, it is a string of math symbols nobody will ever search for.

-58%

less likely to be cited by ChatGPT

+12%

more reactions in the feed

An unmeasured confounder would need to be roughly 4.6 times stronger than anything else in the model to explain this away.

The mechanism is specific, and it is a bug rather than a policy. ChatGPT does not appear to apply Normalization Form Compatibility Decomposition, the standard Unicode procedure that decomposes mathematical alphanumeric characters back to their plain equivalents. One line of code at indexing closes it. Until that ships, unicode-formatted text is functionally invisible to the retriever.

This is not a post problem. It is a string problem.

Here is the part that gets missed, and it is why this is worth a full report rather than a tip. The penalty applies to any page ChatGPT retrieves, not just feed posts. Anywhere those characters sit in the raw text, the words they spell stop being findable. Which means the highest-value real estate you own is often the most heavily formatted.

Your LinkedIn headline. A bolded job title looks sharp in the feed. If it is unicode, the words in it are not queryable.

Your About section. Usually where the actual positioning language lives, and usually formatted.

Executive and founder profiles. OtterlyAI found named individuals account for 91.7 percent of LinkedIn citations against 8.3 percent for company pages and unattributed authors. Your people are the citation surface.

Instagram and Facebook bios. Same characters, same result.

Press releases and profiles pasted from formatted sources. The formatting travels with the paste. Nobody checks.

There is a second cost that predates AI search entirely. Screen readers frequently announce these characters one at a time as mathematical bold capital A, or skip them. LinkedIn's own keyword search does not match them either. So the habit was already costing you accessibility and native search before ChatGPT existed. The citation penalty is the third bill for the same decision.

 

Finding two

The link-in-comments arbitrage

The second penalty is real, but calling it a mistake would be wrong. It is a trade, and once you see both sides of the ledger you may decide to keep making it.

Link-in-comments is the tactic where the post says the real thing is in the comments below. Scrunch found those posts are 31 percent less likely to be cited, and earn about 11 percent more reactions.

The logic is clean. Your post has announced that it is a pointer rather than a source. Generative engines prefer sources that contain the answer, and a post whose main claim is that the content is elsewhere is telling the retriever exactly what it is. Now the other side.

What gets measured Rate Against baseline
The URL posted in the comments 47% About twice a typical source
The LinkedIn post itself 13% About 45% below a typical source
Linked source, retrieved alone 24% Baseline for the comparison
Linked source, retrieved alongside the post 59% A 144% lift

When both the post and the destination appear in the same retrieval set, the destination's citation rate jumps from 24 percent to 59 percent. The post appears to act as corroboration, and it does not need to be cited itself to do that job.

You give up 31 percent of your post's citation odds to buy a 144 percent lift for the thing you actually want cited.

Make that trade when the destination is a page you own and want cited, and the post exists to support it.

Do not make it when the post is the asset. If you have written a genuinely complete, self-contained point, sending readers away costs you the citation and buys a lift for a page that may not deserve it.

The null results

 

What does not predict citation

This is the section to send to whoever asks why the post with four thousand likes did not move anything.

Reaction count. Tested directly, controlling for content, post age, and prompt difficulty. Near-zero predictive power. Hold content constant and a post with 100 reactions gets cited at essentially the same rate as one with 10,000. OtterlyAI reached the same place from a different dataset: likes correlate at Pearson r of -0.06, comments at -0.04, emojis and hashtags at -0.02 each. Not weak positives. Flat to slightly negative.

Length, on its own. Long-form gets cited far more, but the driver is format rather than word count. OtterlyAI measured the raw correlation at Pearson r of 0.03. Cited articles run a median of 1,021 words against 185 for posts, and articles take 72.2 percent of content citations. The lesson is not to write more words. It is to write the kind of thing that has a complete answer in it.

Hashtag volume. Present on 27.4 percent of cited content, correlating at -0.02.

Media. Posts with video averaged 5.69 citations against 6.48 for text-only. Images averaged 5.91. Adding media is a reach decision, not a citation one.

Audience size. Half of all cited URLs in the OtterlyAI set were cited exactly twice, while the single most-cited URL earned 4,685. A small account with the clearest answer beats a large one without it.

Keep going

Studies read properly, including the limitations.

What the research on AI visibility actually says, with the methodology stated and the caveats left in. Written for people who have to defend a number to somebody else.

Subscribe to the newsletter

No spam. Unsubscribe anytime.

 

The signal

 

What does predict citation

Three content dimensions moved the needle, and two of them also help you in the feed.

Attribute ChatGPT citations LinkedIn reactions
Technical detail and specificity +77% No measurable effect
Named entities in plain text +33% +5%
Topic narrowness +18% +13%
Link in comments -31% +11%
Unicode formatting -58% +12%
Reaction count Near zero Not applicable

Technical detail is the dominant finding at 77 percent, and it is not close. This does not mean posting code. It means showing the depth of what you actually know. AI is changing data pipelines is a sentence anyone could write. A sentence naming the three failure modes you have personally watched break a pipeline at scale is a sentence only you could write, and it is matchable to a real question somebody typed.

Named entities work without tagging. ChatGPT reads the text of the post, not the mention graph. Typing the company names in plain text counts. You do not need the at-symbol.

Narrowness beats breadth. Scrunch scored specificity on a one to five scale and the most niche content won on both citations and reactions. This is the one dimension where the two games agree completely.

One more, from the OtterlyAI side: publish under a named human. Named individuals average 8.5 citations per URL against 5.5 for company pages and unattributed authors. Your founder's profile outperforms your brand page, and it is not close.

The caveats

 

The limits of this finding

We would rather you trust the next report than overtrust this one. So here is what this does not say.

It is isolated to ChatGPT. Scrunch ran the same detection across other platforms over the same window. Perplexity showed no statistically significant penalty despite thousands of unicode results in the sample. Google AI Mode and AI Overviews showed zero unicode in snippets across 390,000 observations, and direct embedding tests suggested unicode-bold and plain versions were treated as near-equivalent.

That is not a small caveat. Per OtterlyAI's platform split, ChatGPT accounts for 18.7 percent of LinkedIn citations, behind Perplexity at 43.3 percent and Google AI Overviews at 22.2 percent. The penalty applies to a large and important slice, not the whole pie.

It is time-decaying. The fix is one line of normalization code. The day OpenAI ships it, this finding becomes a historical note. Treat the window as months rather than years.

Correlation is doing some work here. Double machine learning with a full control set is a serious method, and the 4.6x robustness figure is a real check. It is still observational data, not a controlled experiment.

And one honest note about this page. Scrunch found that LinkedIn posts structured around a single question had zero citation payoff and cost 9 percent in reactions. That finding is about the feed. On a web page, an FAQ block with schema markup remains one of the most reliably extractable structures you can publish.

Same tactic, opposite verdict, different surface. Which is the whole argument of this report in one line: know which game the page is playing.

 

This week

 

Six checks to run this week

None of these require budget. Most take under an hour.

  1. 1

    Audit your top five people's headlines and About sections. Paste each into a plain text editor. If the characters change appearance or break, they are unicode. Retype the keyword phrases you want to be found for in ordinary letters.

  2. 2

    Check the last 20 company page posts. Same test. Fix forward rather than editing history, since edits reset post distribution.

  3. 3

    Decide your link-in-comments policy per post, not per account. If the destination is the asset, use it and accept the 31 percent. If the post is the asset, put the link in the body.

  4. 4

    Pull your five best-performing posts and count the named entities. Companies, products, frameworks, people. If the count is zero, you published an opinion where a source should have been.

  5. 5

    Rewrite one broad post as a narrow one. Take your most generic recent piece and rewrite it for the narrowest audience you legitimately serve. Narrowness is the only lever that pays on both scoreboards.

  6. 6

    Scan every surface you paste into. Press releases, client profiles, bios, directory listings, partner pages. Formatted text pasted from a doc carries its characters with it, and nobody checks after the paste.

Common questions

 

The formatting penalty, answered

Why is my content not showing up in ChatGPT?

The two most common formatting causes are unicode text and link-in-comments posting. Scrunch's June 2026 study of 12,000 LinkedIn post observations found unicode-formatted posts were 58 percent less likely to be cited and link-in-comments posts 31 percent less likely. Beyond formatting, the strongest positive predictors were technical specificity at 77 percent, named entities at 33 percent, and topic narrowness at 18 percent.

Does unicode bold text really hurt AI visibility?

On ChatGPT, yes, by a measured 58 percent in the Scrunch dataset. The characters come from the Unicode Mathematical Alphanumeric Symbols block at U+1D400 to U+1D7FF, and ChatGPT does not appear to apply Normalization Form Compatibility Decomposition to convert them back to plain letters. The same test showed no significant penalty on Perplexity and zero unicode in Google AI Mode snippets across 390,000 observations.

Should I stop putting links in the comments?

Only if the post is the asset. The tactic costs the post 31 percent of its citation odds, but when the post and its destination appear together in ChatGPT's retrieval set, the destination's citation rate rises from 24 percent to 59 percent, a 144 percent lift. If you are driving AI attention to a page you own, that trade usually favors you.

Do likes and comments help you get cited by AI?

No. Scrunch found reaction count had near-zero predictive power after controlling for content, post age, and prompt difficulty. Holding content constant, a post with 100 reactions was cited at essentially the same rate as one with 10,000. OtterlyAI's analysis of 1,310,455 citations found likes correlated at Pearson r of -0.06 and comments at -0.04. Engagement measures human reach. Citation measures reference value.

What kind of LinkedIn content gets cited most by AI?

Long-form articles, published by a named person. OtterlyAI found LinkedIn articles took 72.2 percent of content citations at 8.5 citations per URL, against 26.1 percent and 5.9 for posts. Content from named individuals averaged 8.5 citations per URL against 5.5 for company pages. Cited articles ran a median of 1,021 words, though raw word count correlated at only 0.03, meaning the driver is format rather than length.

How long will the unicode penalty last?

Probably not long. The remediation on OpenAI's side is a single normalization call at indexing. Once it ships, the penalty disappears. Treat the advantage as a window measured in months. The underlying principle outlasts it: plain, specific, self-contained text is what generative engines can retrieve, parse, and trust.

Start here

 

See which of the six your own content triggers

You can run these checks yourself with the list above, and you should. Or we will run them across your owned surfaces and show you exactly where you are invisible, which engines it affects, and what a fix is worth. Most teams find the problem in the three places nobody audits: the founder headline, the About section, and press releases pasted in from a formatted document.

Get your AI visibility audit

Sources

Scrunch, "LinkedIn posts that robots can't resist: What the data says about ChatGPT citations," Michael Iannelli, June 5, 2026. 12,000 ChatGPT observations, approximately 4,000 posts analyzed, 21 dimensions, double machine learning. OtterlyAI, "1 in 8 Social Media AI Citations Point to LinkedIn," Rick Tousseyn, June 3, 2026. 1,310,455 LinkedIn AI citations across six platforms, January 1 to June 1, 2026.