Did I get your name right?
@MSGID: 2:5075/21@fidonet fe16e72f
@REPLY: 2:280/5555.1 6a930bff
@PID: SeenBy macOS MVP
@CHRS: UTF-8 4
@TZUTC: 0200
@TID: SeenBy Tosser 0.1
Did I get your name right?
Yes, You did, BTW, can You please check, if my charset is good and
nice? Must be UTF-8, here.
Your message has the "CHRS: UTF-8 4" kludge, so that is correct. I only see ASCII characters though. Which is probably also correct. ASCII is a subset of UTF-8, so labelling an ASCII text as UTF-8 is correct.Thank You sir, but if so, let’s test it further, if You don’t mind:
Thank You sir, but if so, let’s test it further, if You don’t mind:
If everything is configured properly, all of the following characters should be displayed without corruption:
English: The quick brown fox jumps over the lazy dog.
European characters:
Polish: Zażółć gęślą jaźń — ąćęłńóśźż
German: Grüße, Straße, schön — äöüÄÖÜß
French: déjà vu, café, naïve, œuvre — àâçéèêëîïôùûüÿ Spanish: ¿Cómo está? ¡Muy bien! — áéíóúüñ
Czech: Příliš žluťoučký kůň úpěl ďábelské ódy.
Cyrillic:
Russian: Съешь ещё этих мягких французских булок, да выпей чаю.
Ukrainian: Ґанок, їжак, пір’я, єдність — Ґґ Єє Іі Її.
Other scripts:
Greek: Ελληνικά — Καλημέρα κόσμε.
Hebrew: עברית — שלום עולם.
Arabic: العربية — مرحبا بالعالم.
Chinese: 中文 — 你好,世界。
Japanese: 日本語 — こんにちは世界。
Korean: 한국어 — 안녕하세요 세계.
Typography:
“Double quotes” ‘single quotes’ — en dash – em dash … ellipsis
© ® ™ § ¶ † ‡ • → ← ↑ ↓ ↔
½ ¼ ¾ ± × ÷ ≠ ≤ ≥ ≈ ∞ √ ∑ π
Currencies:
€ £ ¥ ₽ ₴ ₩ ₹ $ ¢
Emoji:
😀 😎 🚀 ❤️ 👍 🔥 🐈 🌍 ✈️ ☕️
Mixed UTF-8 test:
Wrocław → Москва → Αθήνα → ירושלים → 北京 → 東京 → 서울 🚀
A particularly useful corruption test:
Zażółć gęślą jaźń / Привет, мир! / €100 / café / 日本語 / 😀
If UTF-8 is being interpreted incorrectly as Windows-1252 or
ISO-8859-1, this line will usually make the problem immediately
obvious.
With best regards,
So... How about enticing your NC and RC to participate in the UTF-8 nodelist project?
Hello Michiel van der Vlist!
Did I get your name right?
Yes, You did, BTW, can You please check, if my charset is good and
nice? Must be UTF-8, here.
--- SeenBy 0.1
* Origin: SeenBy macOS MVP (2:5075/21@fidonet)
But he wrote кириллица wrong. It is with _one_ р.That’s what was I seening and it's a big surprise to me, the name Eugene must contain P, are You sure about that?
Hello Евгений,
Thank You sir, but if so, let’s test it further, if You don’t mind:
If everything is configured properly, all of the following characters should be displayed without corruption:
Thank you — this is actually a very useful result.
There is one particularly interesting observation: in your quoted copy
of my test message, *all* the original characters are displayed
correctly on my side, including Hebrew, Arabic, Chinese, Japanese,
Korean, the currency symbols and emoji which you see as squares.
This seems to prove that, at the transport level, everything is
working correctly. Those characters have made the complete round trip:
me → Fidonet → you → quoted reply → Fidonet → me
without being damaged or substituted.
In other words, the actual Unicode code points appear to survive
perfectly well. What you are seeing as squares is therefore very
likely a rendering problem rather than an UTF-8 encoding problem — probably missing glyphs in the font used by your reader, or the reader
not performing font fallback.
Could you please try one more small experiment?
Copy one of the squares corresponding to, say, the Japanese 日,
Chinese 中, or Hebrew ש from my original message and paste it into another Unicode-aware application — a browser, a reasonably modern
text editor, Word, etc.
If the proper character suddenly appears there instead of the square,
I think we can consider the diagnosis conclusive: the character is
present and correctly decoded, but your Fidonet reader simply cannot render its glyph.
It would be even more interesting if an editor capable of showing the Unicode code point reports U+65E5 for 日, U+4E2D for 中, etc.
So... How about enticing your NC and RC to participate in the
UTF-8 nodelist project?
As for my NC — amusingly enough, you already know him from this very discussion. It is Eugene Subbotin, 2:5075/35, the same guy with whom
this whole correspondence started. And he is already quite actively involved in the subject, so I don't think much enticing will be
necessary. :)
The RC is another matter, but considering that we have just
successfully sent Japanese and emoji from my system to yours and back again — through Fidonet — without losing a single code point, I think we now have a rather nice practical demonstration to show him.
Did I get your name right?
Yes, You did, BTW, can You please check, if my charset is good andBut he wrote кириллица wrong. It is with _one_ р.
nice? Must be UTF-8, here.
No surprise. You entered it on your system, so your system can display them.
It is what I expected. The transport chain and the message base is not encoding aware. That it is fully 8 bit transparent is what we have known for decades.
The only one I have at hand here at my point system is notepad. Almost all of it seems to be displayed correct in notepad, except for the emois. I only see a heart and an airplane. The rest is squares as well in notepad.
RC50 once told me that if we could find five NCs interested he was willing to go for it. IIRC that was about ten years ago....
So let's find four more R50 NCs. ...
Hello Chris Jacobs!
But he wrote кириллица wrong. It is with _one_ р.
Hello Евгений,That’s what was I seening and it's a big surprise to me, the name
Eugene must contain P, are You sure about that?
--- SeenBy 0.1
* Origin: SeenBy macOS MVP (2:5075/21@fidonet)
Works for me. Everything is displayed correctly, including the emojis.
The name Eugene must neither contain a p nor a р. I was refering to the word "Cyrrilic".
Works for me. Everything is displayed correctly, including the emojis.As it should in a modern software. Thank You for the test!
BTW, are You using Your own solution?
That߷�s what was I seening^^^
and it's a big surprise to me, the name^
Emoji:
😀 😎 🚀 ❤️ 👍 🔥 🐈 🌍 ✈️ ☕️
Yes, You did, BTW, can You please check, if my charset is good
and nice? Must be UTF-8, here.
Your message suddenly brought to light a huge problem. And the problem isnʼt with the reader, but with the tosser. The thing is, the original message with a cyrillic UTF-8 name in the To: field, for some reason, didn’t make it into my JAM database for this echo.
However, it passed through my station and reached 2:5075/21 without
any issues. Thereʼs clearly a problem with the code of the hpt tosser from the Husky project: Iʼve already run into this once before—it has issues with cyrillic names in message headers, even in CP866. All in
all, this will require further study of the issue.
This shows that the technical readiness to use non-ASCII names in the “From” and “To” fields is still limited and can cause problems even
with fairly modern software.
Not only did my tosser, for some reason, fail to store a record with
the To-field "������� �㡡�⨭" in the JAM database - and the reasons
for this are still unknown
Following up on that point, a few more thoughts came to mind.
Not only did my tosser, for some reason, fail to store a record with the To-field “Евгений Субботин” in the JAM database — and the reasons for
this are still unknown — but there may be even more problems overall.
BTW, I had to manually enter your cyrillic name in the header. For
some reason copy/paste does not work for the header....
BTW, I had to manually enter your cyrillic name in the header.
For some reason copy/paste does not work for the header....
Alt-P only works for the body. In the header you have to right-click
on your mouse (or press Alt+Space -> Edit -> Paste)
Not only did my tosser, for some reason, fail to store a record
with the To-field “Евгений Субботин” in the JAM database — and the
reasons for this are still unknown — but there may be even more
problems overall.
For the record, I'm using the same (latest) version of hpt for Linux,
and it seems that message stored in my JAM database just fine. I can
grep the name right out of the .jhr file. Unless I'm not understanding this correctly..
Alt-P only works for the body. In the header you have to
right-click on your mouse (or press Alt+Space -> Edit -> Paste)
I know and IIRC that worked before.
But... It works with ASCII text. Not with the Cyrillic name I tried to copy/paste.
Odd...
Can you copy/paste cyrrylics into the header?
Your message suddenly brought to light a huge problem. And the
problem isnʼt with the reader, but with the tosser. The thing is,
the original message with a cyrillic UTF-8 name in the To: field,
for some reason, didn’t make it into my JAM database for this
echo.
However, it passed through my station and reached 2:5075/21
without any issues. Thereʼs clearly a problem with the code of
the hpt tosser from the Husky project: Iʼve already run into this
once before—it has issues with cyrillic names in message headers,
even in CP866. All in all, this will require further study of the
issue.
This shows that the technical readiness to use non-ASCII names in
the “From” and “To” fields is still limited and can cause
problems even with fairly modern software.
Following up on that point, a few more thoughts came to mind.
Not only did my tosser, for some reason, fail to store a record with
the To-field “Евгений Субботин” in the JAM database — and the reasons
for this are still unknown — but there may be even more problems overall.
For example, the name “Александр Христофоров” in UTF-8 will take up 41
bytes, which won’t fit within the 36 bytes defined in FTS-0001.
For example, FSP-1030 proposed using the kludges ^AUCSFROM:, ^AUCSTO:,
and ^AUCSSUBJ: to include UTF-8 header fields in messages, but this
was not adopted as a standard.
This means that five bytes will be truncated along the way, which will corrupt the UTF-8 text. With GoldED+, exporting truncation will occur correctly — by character rather than by byte — but it may work differently on other systems.
So there are far more problems, and they require a technical solution
and standardization first before we can use such fields from the UTF-8 nodelist in headers.
For example, the name “Александр Христофоров” in UTF-8 will take
up 41 bytes, which won’t fit within the 36 bytes defined in
FTS-0001.
| Sysop: | DaiTengu |
|---|---|
| Location: | Appleton, WI |
| Users: | 1,136 |
| Nodes: | 10 (0 / 10) |
| Uptime: | 08:34:22 |
| Calls: | 14,591 |
| Calls today: | 2 |
| Files: | 186,473 |
| D/L today: |
4,886 files (1,449M bytes) |
| Messages: | 2,577,398 |