Sign inGet Started

SMS smart encoding

Typographic substitutions can push an SMS out of GSM-7. A word processor turns ' into ’, pasted text contains an em dash, or a form field carries a non-breaking space. Each looks similar to a GSM-7 character but one non-GSM character converts the whole body to UCS-2, reducing the capacity of each segment.

Smart encoding replaces those characters with their GSM-7 equivalents before Bird hands the message to the carrier, so a body that only looked non-GSM stays GSM-7 and bills as fewer segments.

Turning it on

It is off by default, because it changes the text your recipient reads. Set options.smart_encoding to true on a send:

await bird.sms.send({
  to: "+31612345678",
  from: "Bird",
  text: "Your order shipped, track it here…",
  category: "transactional",
  options: { smart_encoding: true },
});

The Go and PHP SDKs take it as a nullable field, so leaving it unset takes Bird's default rather than sending false.

Set smart encoding per message. Bird does not provide an account-wide or per-sender setting.

When Bird keeps UCS-2

Replacement applies only when it converts the whole body to GSM-7. If any character remains outside GSM-7, Bird sends the body exactly as supplied.

A single emoji, CJK character, or accented letter outside GSM-7 keeps the whole message in UCS-2:

Code example
{ "text": "Hoe gaat het vandaag? á “" }

The “ is replaceable, but á is not in GSM-7 and has no entry in the table, so this body sends unchanged as UCS-2. GSM-7 includes à, ä, å, é, è, ì, ò, ù, ö, ñ, ü but not á, í, or ú.

Accented letters are deliberately absent from the table. Folding á to a would change the word rather than approximate it, which is the wrong trade for a sender writing Spanish or Portuguese.

Normalization affects combining accents. In NFD text, a decomposed ê stores the accent as a separate combining mark, so smart encoding converts it to e^. A precomposed ê remains unchanged and uses UCS-2.

What it does not do

  • It never truncates. A body over the 12-segment cap is rejected with a 422, and the cap is measured after replacement, so opting in can bring an oversized body under it.
  • It does not touch a body that is already GSM-7, which has no segments to save.
  • It does not change what you are billed for beyond the segment count: the count is computed from the body as sent.

Reading back what happened

The message reports both halves. text is the body as sent, and options.smart_encoding reports the setting that applied:

Code example
{
  "text": "Your order shipped - track it here...",
  "segments": { "count": 1, "encoding": "GSM_7BIT", "characters": 36 },
  "options": { "smart_encoding": true }
}

An encoding of UCS2 on a send you opted in tells you the body could not be brought into GSM-7. To see what would happen to a specific message before you send it, paste the text into the SMS segment calculator.

The full replacement table

The replacement table lists every character changed when smart encoding is on. Other characters, including those already in GSM-7, remain unchanged.

ReplacementCharacters replaced
spaceU+00A0, U+2000, U+2001, U+2002, U+2003, U+2004, U+2005, U+2006, U+2007, U+2008, U+2009, U+200A, U+202F, U+205F, U+3000
!ǃ U+01C3, ︕ U+FE15, ﹗ U+FE57, ! U+FF01
!!‼ U+203C
"« U+00AB, » U+00BB, ʺ U+02BA, ˮ U+02EE, “ U+201C, ” U+201D, „ U+201E, ‟ U+201F, ❝ U+275D, ❞ U+275E, 〝 U+301D, 〞 U+301E, " U+FF02
#﹟ U+FE5F, # U+FF03
$﹩ U+FE69, $ U+FF04
%﹪ U+FE6A, % U+FF05
&﹠ U+FE60, & U+FF06
'` U+0060, ´ U+00B4, ʹ U+02B9, ʻ U+02BB, ʼ U+02BC, ʽ U+02BD, ˈ U+02C8, ˊ U+02CA, ˋ U+02CB, U+0313, U+0314, ‘ U+2018, ’ U+2019, ‛ U+201B, ❛ U+275B, ❜ U+275C, ︐ U+FE10, ︑ U+FE11, ' U+FF07
(❨ U+2768, ❪ U+276A, ⟮ U+27EE, ⦅ U+2985, ﹙ U+FE59, ( U+FF08
)❩ U+2769, ❫ U+276B, ⟯ U+27EF, ⦆ U+2986, ﹚ U+FE5A, ) U+FF09
*⁎ U+204E, ∗ U+2217, ⊛ U+229B, ✢ U+2722, ✣ U+2723, ✤ U+2724, ✥ U+2725, ✱ U+2731, ✲ U+2732, ✳ U+2733, ✺ U+273A, ✻ U+273B, ✼ U+273C, ✽ U+273D, ❃ U+2743, ❉ U+2749, ❊ U+274A, ❋ U+274B, ⧆ U+29C6, ﹡ U+FE61, * U+FF0A
+˖ U+02D6, ﹢ U+FE62, + U+FF0B
,U+0326, ‚ U+201A, 、 U+3001, ﹐ U+FE50, ﹑ U+FE51, , U+FF0C, 、 U+FF64
-‐ U+2010, – U+2013, — U+2014, ― U+2015, • U+2022, ⁃ U+2043, ⎼ U+23BC, ⎽ U+23BD, ﹣ U+FE63, - U+FF0D
.。 U+3002, ﹒ U+FE52, . U+FF0E, 。 U+FF61
...… U+2026
/÷ U+00F7, U+0337, U+0338, ⁄ U+2044, ∕ U+2215, ⧸ U+29F8, / U+FF0F
00 U+FF10
11 U+FF11
1/2½ U+00BD
1/4¼ U+00BC
22 U+FF12
33 U+FF13
3/4¾ U+00BE
44 U+FF14
55 U+FF15
66 U+FF16
77 U+FF17
88 U+FF18
99 U+FF19
:ː U+02D0, ˸ U+02F8, ⦂ U+2982, ꞉ U+A789, ︓ U+FE13, : U+FF1A
;⁏ U+204F, ︔ U+FE14, ﹔ U+FE54, ; U+FF1B
<‹ U+2039, ﹤ U+FE64, < U+FF1C
=U+0347, ꞊ U+A78A, ﹦ U+FE66, = U+FF1D
>› U+203A, ﹥ U+FE65, > U+FF1E
?︖ U+FE16, ﹖ U+FE56, ? U+FF1F
@﹫ U+FE6B, @ U+FF20
Aᴀ U+1D00, A U+FF21
Bʙ U+0299, B U+FF22
Cᴄ U+1D04, C U+FF23
Dᴅ U+1D05, D U+FF24
Eᴇ U+1D07, E U+FF25
Fꜰ U+A730, F U+FF26
Gɢ U+0262, G U+FF27
Hʜ U+029C, H U+FF28
Iɪ U+026A, I U+FF29
Jᴊ U+1D0A, J U+FF2A
Kᴋ U+1D0B, K U+FF2B
Lʟ U+029F, L U+FF2C
Mᴍ U+1D0D, M U+FF2D
Nɴ U+0274, N U+FF2E
Oᴏ U+1D0F, O U+FF2F
Pᴘ U+1D18, P U+FF30
QQ U+FF31
Rʀ U+0280, R U+FF32
Sꜱ U+A731, S U+FF33
Tᴛ U+1D1B, T U+FF34
Uᴜ U+1D1C, U U+FF35
Vᴠ U+1D20, V U+FF36
Wᴡ U+1D21, W U+FF37
XX U+FF38
Yʏ U+028F, Y U+FF39
Zᴢ U+1D22, Z U+FF3A
[[ U+FF3B
\U+20E5, ⧵ U+29F5, ⧹ U+29F9, ﹨ U+FE68, \ U+FF3C
]] U+FF3D
^ˆ U+02C6, U+0302, U+1DCD, ^ U+FF3E
_U+0332, ‗ U+2017, _ U+FF3F
{❴ U+2774, ﹛ U+FE5B, { U+FF5B
|U+20D2, U+20D3, ∣ U+2223, ⎜ U+239C, ⎟ U+239F, ⎸ U+23B8, ⎹ U+23B9, ⏐ U+23D0, | U+FF5C
}❵ U+2775, ﹜ U+FE5C, } U+FF5D
~˜ U+02DC, ˷ U+02F7, U+0303, U+0330, U+0334, ∼ U+223C, ~ U+FF5E
removedU+200B, U+2028, U+2029, U+2060, U+FEFF

Whitespace variants collapse to a single space (U+0020). The final row lists zero-width and formatting characters that are deleted rather than replaced.

Next steps