• Name in Cyrrilic

    From Michiel van der Vlist@2:280/5555.1 to Евгений Субботин on Sat Aug 29 18:38:11 2026
    Hello Евгений,

    Did I get your name right?

    Cheers, Michiel

    --- GoldED+/W32-MINGW 1.1.5-b20260829
    * Origin: Klein Schnøørd (2:280/5555.1)
  • From Dmitry Lipatnikov@2:5075/21 to Michiel van der Vlist on Sat Aug 29 20:12:55 2026

    Hello Michiel van der Vlist!

    Did I get your name right?

    Yes, You did, BTW, can You please check, if my charset is good and nice? Must be UTF-8, here.

    --- SeenBy 0.1
    * Origin: SeenBy macOS MVP (2:5075/21@fidonet)
  • From Michiel van der Vlist@2:280/5555.1 to Dmitry Lipatnikov on Sat Aug 29 20:20:13 2026
    Hello Dmitry,

    On 29 Aug 26 20:12, you wrote to me:

    @MSGID: 2:5075/21@fidonet fe16e72f
    @REPLY: 2:280/5555.1 6a930bff
    @PID: SeenBy macOS MVP
    @CHRS: UTF-8 4
    @TZUTC: 0200
    @TID: SeenBy Tosser 0.1

    Did I get your name right?

    Yes, You did, BTW, can You please check, if my charset is good and
    nice? Must be UTF-8, here.

    Your message has the "CHRS: UTF-8 4" kludge, so that is correct. I only see ASCII characters though. Which is probably also correct. ASCII is a subset of UTF-8, so labelling an ASCII text as UTF-8 is correct.

    Some old software may not recognise the kludge and so just assume the default Fidonet character encoding. Which is ASCII, so all is well. ;-)

    Cheers, Michiel

    --- GoldED+/W32-MINGW 1.1.5-b20260829
    * Origin: Klein Schnøørd (2:280/5555.1)
  • From Dmitry Lipatnikov@2:5075/21 to Michiel van der Vlist on Sat Aug 29 20:34:14 2026

    Hello Michiel van der Vlist!

    Your message has the "CHRS: UTF-8 4" kludge, so that is correct. I only see ASCII characters though. Which is probably also correct. ASCII is a subset of UTF-8, so labelling an ASCII text as UTF-8 is correct.
    Thank You sir, but if so, let’s test it further, if You don’t mind:


    If everything is configured properly, all of the following characters should be displayed without corruption:

    English: The quick brown fox jumps over the lazy dog.

    European characters:
    Polish: Zażółć gęślą jaźń — ąćęłńóśźż
    German: Grüße, Straße, schön — äöüÄÖÜß
    French: déjà vu, café, naïve, œuvre — àâçéèêëîïôùûüÿ Spanish: ¿Cómo está? ¡Muy bien! — áéíóúüñ
    Czech: Příliš žluťoučký kůň úpěl ďábelské ódy.

    Cyrillic:
    Russian: Съешь ещё этих мягких французских булок, да выпей чаю.
    Ukrainian: Ґанок, їжак, пір’я, єдність — Ґґ Єє Іі Її.

    Other scripts:
    Greek: Ελληνικά — Καλημέρα κόσμε.
    Hebrew: עברית — שלום עולם.
    Arabic: العربية — مرحبا بالعالم.
    Chinese: 中文 — 你好,世界。
    Japanese: 日本語 — こんにちは世界。
    Korean: 한국어 — 안녕하세요 세계.

    Typography:
    “Double quotes” ‘single quotes’ — en dash – em dash … ellipsis
    © ® ™ § ¶ † ‡ • → ← ↑ ↓ ↔
    ½ ¼ ¾ ± × ÷ ≠ ≤ ≥ ≈ ∞ √ ∑ π

    Currencies:
    € £ ¥ ₽ ₴ ₩ ₹ $ ¢

    Emoji:
    😀 😎 🚀 ❤️ 👍 🔥 🐈 🌍 ✈️ ☕️

    Mixed UTF-8 test:
    Wrocław → Москва → Αθήνα → ירושלים → 北京 → 東京 → 서울 🚀

    A particularly useful corruption test:
    Zażółć gęślą jaźń / Привет, мир! / €100 / café / 日本語 / 😀

    If UTF-8 is being interpreted incorrectly as Windows-1252 or ISO-8859-1, this line will usually make the problem immediately obvious.


    With best regards,
    --- SeenBy 0.1
    * Origin: SeenBy macOS MVP (2:5075/21@fidonet)
  • From Michiel van der Vlist@2:280/5555.1 to Dmitry Lipatnikov on Sat Aug 29 21:26:12 2026
    Hello Dmitry,

    On 29 Aug 26 20:34, you wrote to me:

    Thank You sir, but if so, let’s test it further, if You don’t mind:

    I don't mind...

    If everything is configured properly, all of the following characters should be displayed without corruption:

    English: The quick brown fox jumps over the lazy dog.

    OK

    Zeven schele geitjes met scheefgegroeide gewrichten gingen gezellig samen googelen in Scheveningen.

    European characters:
    Polish: Zażółć gęślą jaźń — ąćęłńóśźż
    German: Grüße, Straße, schön — äöüÄÖÜß
    French: déjà vu, café, naïve, œuvre — àâçéèêëîïôùûüÿ Spanish: ¿Cómo está? ¡Muy bien! — áéíóúüñ
    Czech: Příliš žluťoučký kůň úpěl ďábelské ódy.

    Cyrillic:
    Russian: Съешь ещё этих мягких французских булок, да выпей чаю.
    Ukrainian: Ґанок, їжак, пір’я, єдність — Ґґ Єє Іі Її.

    Looks OK as well. How wonderfull that all of this can be displayed in one and the same message. That is impossible with a one byte character encoding! Vivat UTF-8!

    Other scripts:
    Greek: Ελληνικά — Καλημέρα κόσμε.
    Hebrew: עברית — שלום עולם.
    Arabic: العربية — مرحبا بالعالم.
    Chinese: 中文 — 你好,世界。
    Japanese: 日本語 — こんにちは世界。
    Korean: 한국어 — 안녕하세요 세계.

    Greek looks OK, but the others are displayed as squares. No surprise. Apparently my version of Windows (7 professional) does not have the glyphs for those. It does not mean that anything is misconfigured on the Fidonet level.

    Typography:
    “Double quotes” ‘single quotes’ — en dash – em dash … ellipsis
    © ® ™ § ¶ † ‡ • → ← ↑ ↓ ↔
    ½ ¼ ¾ ± × ÷ ≠ ≤ ≥ ≈ ∞ √ ∑ π

    Looks OK.

    Currencies:
    € £ ¥ ₽ ₴ ₩ ₹ $ ¢

    The Euro, Pound, Yen, Dollar and cent are displayed correctly. The four between Yen and Dollar are displayed as squares.

    Emoji:
    😀 😎 🚀 ❤️ 👍 🔥 🐈 🌍 ✈️ ☕️

    All squares.

    Mixed UTF-8 test:
    Wrocław → Москва → Αθήνα → ירושלים → 北京 → 東京 → 서울 🚀

    First three OK, the rest squares.

    A particularly useful corruption test:
    Zażółć gęślą jaźń / Привет, мир! / €100 / café / 日本語 / 😀

    OK till "café", last two squares

    If UTF-8 is being interpreted incorrectly as Windows-1252 or
    ISO-8859-1, this line will usually make the problem immediately
    obvious.

    That is not the case.

    With best regards,

    So... How about enticing your NC and RC to participate in the UTF-8 nodelist project?


    Cheers, Michiel

    --- GoldED+/W32-MINGW 1.1.5-b20260829
    * Origin: Klein Schnøørd (2:280/5555.1)
  • From Dmitry Lipatnikov@2:5075/21 to Michiel van der Vlist on Sat Aug 29 21:50:02 2026

    Hello Michiel van der Vlist!


    Thank you — this is actually a very useful result.

    There is one particularly interesting observation: in your quoted copy of my test message, *all* the original characters are displayed correctly on my side, including Hebrew, Arabic, Chinese, Japanese, Korean, the currency symbols and emoji which you see as squares.

    This seems to prove that, at the transport level, everything is working correctly. Those characters have made the complete round trip:

    me → Fidonet → you → quoted reply → Fidonet → me

    without being damaged or substituted.

    In other words, the actual Unicode code points appear to survive perfectly well. What you are seeing as squares is therefore very likely a rendering problem rather than an UTF-8 encoding problem — probably missing glyphs in the font used by your reader, or the reader not performing font fallback.

    Could you please try one more small experiment?

    Copy one of the squares corresponding to, say, the Japanese 日, Chinese 中, or Hebrew ש from my original message and paste it into another Unicode-aware application — a browser, a reasonably modern text editor, Word, etc.

    If the proper character suddenly appears there instead of the square, I think we can consider the diagnosis conclusive: the character is present and correctly decoded, but your Fidonet reader simply cannot render its glyph.

    It would be even more interesting if an editor capable of showing the Unicode code point reports U+65E5 for 日, U+4E2D for 中, etc.

    So... How about enticing your NC and RC to participate in the UTF-8 nodelist project?

    As for my NC — amusingly enough, you already know him from this very discussion. It is Eugene Subbotin, 2:5075/35, the same guy with whom this whole correspondence started. And he is already quite actively involved in the subject, so I don't think much enticing will be necessary. :)

    The RC is another matter, but considering that we have just successfully sent Japanese and emoji from my system to yours and back again — through Fidonet — without losing a single code point, I think we now have a rather nice practical demonstration to show him.

    Vivat UTF-8 indeed. :)

    --- SeenBy 0.1
    * Origin: SeenBy macOS MVP (2:5075/21@fidonet)
  • From Chris Jacobs@2:280/5555.11 to Dmitry Lipatnikov on Sat Aug 29 21:52:01 2026

    Hello Dmitry!

    29 Aug 26 20:12, you wrote to Michiel van der Vlist:


    Hello Michiel van der Vlist!

    Did I get your name right?

    Yes, You did, BTW, can You please check, if my charset is good and
    nice? Must be UTF-8, here.

    But he wrote кириллица wrong. It is with _one_ р.

    --- SeenBy 0.1
    * Origin: SeenBy macOS MVP (2:5075/21@fidonet)

    Chris


    --- Binkd, FMail, Golded+
    * Origin: https://drschrisjacobs.nl (2:280/5555.11)
  • From Dmitry Lipatnikov@2:5075/21 to Michiel van der Vlist on Sat Aug 29 22:08:44 2026

    Hello Chris Jacobs!

    But he wrote кириллица wrong. It is with _one_ р.

    Hello Евгений,
    That’s what was I seening and it's a big surprise to me, the name Eugene must contain P, are You sure about that?

    --- SeenBy 0.1
    * Origin: SeenBy macOS MVP (2:5075/21@fidonet)
  • From Yegor Gluhov@2:382/736 to Dmitry Lipatnikov on Sat Aug 29 22:06:17 2026
    Hello Dmitry!

    29 Aug 26 20:34, you wrote to Michiel van der Vlist:

    Thank You sir, but if so, let’s test it further, if You don’t mind:

    If everything is configured properly, all of the following characters should be displayed without corruption:

    Works for me. Everything is displayed correctly, including the emojis.

    Yegor
    --- AmberEdit/linux 0.5.3
    * Origin: to err is human, but to really f.. things up you need AI (2:382/736)
  • From Michiel van der Vlist@2:280/5555.1 to Dmitry Lipatnikov on Sat Aug 29 22:02:30 2026
    Hello Dmitry,

    On 29 Aug 26 21:50, you wrote to me:

    Thank you — this is actually a very useful result.

    You'r welcome.

    There is one particularly interesting observation: in your quoted copy
    of my test message, *all* the original characters are displayed
    correctly on my side, including Hebrew, Arabic, Chinese, Japanese,
    Korean, the currency symbols and emoji which you see as squares.

    No surprise. You entered it on your system, so your system can display them.

    This seems to prove that, at the transport level, everything is
    working correctly. Those characters have made the complete round trip:

    me → Fidonet → you → quoted reply → Fidonet → me

    without being damaged or substituted.

    It is what I expected. The transport chain and the message base is not encoding aware. That it is fully 8 bit transparent is what we have known for decades.

    In other words, the actual Unicode code points appear to survive
    perfectly well. What you are seeing as squares is therefore very
    likely a rendering problem rather than an UTF-8 encoding problem — probably missing glyphs in the font used by your reader, or the reader
    not performing font fallback.

    Yes, it must be something like that. It does not bother me all that much. The examples are quit exotic and it is highly unlikely that these missing characters will appear in an actual "non test" Fidonet message.

    Could you please try one more small experiment?

    Copy one of the squares corresponding to, say, the Japanese 日,
    Chinese 中, or Hebrew ש from my original message and paste it into another Unicode-aware application — a browser, a reasonably modern
    text editor, Word, etc.

    The only one I have at hand here at my point system is notepad. Almost all of it seems to be displayed correct in notepad, except for the emois. I only see a heart and an airplane. The rest is squares as well in notepad.

    If the proper character suddenly appears there instead of the square,
    I think we can consider the diagnosis conclusive: the character is
    present and correctly decoded, but your Fidonet reader simply cannot render its glyph.

    Yep.

    It would be even more interesting if an editor capable of showing the Unicode code point reports U+65E5 for 日, U+4E2D for 中, etc.

    In notepad, they are displayed as two symbols each which are meaningless to me.

    So... How about enticing your NC and RC to participate in the
    UTF-8 nodelist project?

    As for my NC — amusingly enough, you already know him from this very discussion. It is Eugene Subbotin, 2:5075/35, the same guy with whom
    this whole correspondence started. And he is already quite actively involved in the subject, so I don't think much enticing will be
    necessary. :)

    ;-)

    The RC is another matter, but considering that we have just
    successfully sent Japanese and emoji from my system to yours and back again — through Fidonet — without losing a single code point, I think we now have a rather nice practical demonstration to show him.

    RC50 once told me that if we could find five NCs interested he was willing to go for it. IIRC that was about ten years ago....

    So let's find four more R50 NCs. ...


    Cheers, Michiel

    --- GoldED+/W32-MINGW 1.1.5-b20260829
    * Origin: Klein Schnøørd (2:280/5555.1)
  • From Yegor Gluhov@2:382/736 to Chris Jacobs on Sat Aug 29 22:16:16 2026
    Hello Chris!

    29 Aug 26 21:52, you wrote to Dmitry Lipatnikov:

    Did I get your name right?
    Yes, You did, BTW, can You please check, if my charset is good and
    nice? Must be UTF-8, here.
    But he wrote кириллица wrong. It is with _one_ р.

    Even with one l, when it's ћирилица. :-)

    Yegor
    --- AmberEdit/linux 0.5.3
    * Origin: to err is human, but to really f.. things up you need AI (2:382/736)
  • From Dmitry Lipatnikov@2:5075/21 to Michiel van der Vlist on Sat Aug 29 22:41:14 2026

    Hello Michiel van der Vlist!

    No surprise. You entered it on your system, so your system can display them.

    Yeah, and since my reader uses latest up-to-date MacOS libraries involved in rendering, they are looking correctly.

    It is what I expected. The transport chain and the message base is not encoding aware. That it is fully 8 bit transparent is what we have known for decades.

    Well shit happens, a tosser of sorts, a lot of things can spoil the picture here and there, I don't know if You remember, there were a lot of UUCP systems cutting 8th bit and using 7 bit encoding to save some space. That's what made KOI-8r popular, BTW. If You cut it that way, You’ll have transliteration in latin letter instead of Cyrillic.

    The only one I have at hand here at my point system is notepad. Almost all of it seems to be displayed correct in notepad, except for the emois. I only see a heart and an airplane. The rest is squares as well in notepad.

    Perhaps, European-oriented builds of some software may only be focused on European fonts, and having issues with eastern symbols, which were not a goual at all.


    RC50 once told me that if we could find five NCs interested he was willing to go for it. IIRC that was about ten years ago....

    So let's find four more R50 NCs. ...

    This is also a step to this, showing there is a software readiness already happened ofcourse there would be people who are still using CP/M based software (I heard about it in couple of times) but yep, we are ready. At least my software indeed works with UTF-8 inside, making conversion during tossing process, using cludges to detect direction.


    --- SeenBy 0.1
    * Origin: SeenBy macOS MVP (2:5075/21@fidonet)
  • From Chris Jacobs@2:280/5555.11 to Dmitry Lipatnikov on Sat Aug 29 22:39:37 2026

    Hello Dmitry!

    29 Aug 26 22:08, you wrote:


    Hello Chris Jacobs!

    But he wrote кириллица wrong. It is with _one_ р.

    Hello Евгений,
    That’s what was I seening and it's a big surprise to me, the name
    Eugene must contain P, are You sure about that?

    The name Eugene must neither contain a p nor a р. I was refering to the word "Cyrrilic".

    --- SeenBy 0.1
    * Origin: SeenBy macOS MVP (2:5075/21@fidonet)

    Chris


    --- Binkd, FMail, Golded+
    * Origin: https://drschrisjacobs.nl (2:280/5555.11)
  • From Dmitry Lipatnikov@2:5075/21 to Yegor Gluhov on Sat Aug 29 22:48:07 2026

    Hello Yegor Gluhov!

    Works for me. Everything is displayed correctly, including the emojis.

    As it should in a modern software. Thank You for the test!
    BTW, are You using Your own solution?

    --- SeenBy 0.1
    * Origin: SeenBy macOS MVP (2:5075/21@fidonet)
  • From Dmitry Lipatnikov@2:5075/21 to Michiel van der Vlist on Sat Aug 29 23:02:31 2026

    Hello Chris Jacobs!

    The name Eugene must neither contain a p nor a р. I was refering to the word "Cyrrilic".

    Sorry, my bad, didn’t noticed that word at all, since it was looking natural to me, and typo was not that important to notice in a context.

    --- SeenBy 0.1
    * Origin: SeenBy macOS MVP (2:5075/21@fidonet)
  • From Yegor Gluhov@2:382/736 to Dmitry Lipatnikov on Sat Aug 29 23:01:43 2026
    Hello Dmitry!

    29 Aug 26 22:48, you wrote to me:

    Works for me. Everything is displayed correctly, including the emojis.
    As it should in a modern software. Thank You for the test!
    BTW, are You using Your own solution?

    Yeah, this one: https://github.com/jegornet/amberedit

    Yegor
    --- AmberEdit/linux 0.5.3
    * Origin: to err is human, but to really f.. things up you need AI (2:382/736)
  • From Wilfred van Velzen@2:280/464 to Dmitry Lipatnikov on Sat Aug 29 23:36:23 2026
    Hi Dmitry,

    On 2026-08-29 22:08:44, you wrote to Michiel van der Vlist:

    That߷�s what was I seening
    ^^^
    Why is there an utf-8 character for the single quote character.

    and it's a big surprise to me, the name
    ^
    And here it is the simple ascii version of the single quote.

    The first one is annoying if you can't display (most) utf-8 characters...


    Bye, Wilfred.

    --- FMail-lnx64 2.3.4.1-B20260520
    * Origin: FMail development HQ (2:280/464)
  • From Eugene Subbotin@2:5075/35 to Michiel van der Vlist on Sun Aug 30 00:48:08 2026
    Hello Michiel!

    Saturday August 29 2026 21:26, you wrote to Dmitry Lipatnikov:

    Emoji:
    😀 😎 🚀 ❤️ 👍 🔥 🐈 🌍 ✈️ ☕️

    MvdV> All squares.

    This is to be expected, since the version of Unicode in your Windows 7 didn't yet support emojis, so neither the terminal nor the fonts support them.

    Eugene

    ... It's full of stars!
    --- GoldED+/BSD 1.1.5-b20260828 (NetBSD 11.0 Intel Core Haswell)
    * Origin: FireFox Station (2:5075/35)
  • From Eugene Subbotin@2:5075/35 to Michiel van der Vlist on Sun Aug 30 03:38:44 2026
    Hello Michiel!

    Saturday August 29 2026 20:20, you wrote to Dmitry Lipatnikov:

    Yes, You did, BTW, can You please check, if my charset is good
    and nice? Must be UTF-8, here.

    MvdV> Your message has the "CHRS: UTF-8 4" kludge, so that is correct. I
    MvdV> only see ASCII characters though. Which is probably also correct.
    MvdV> ASCII is a subset of UTF-8, so labelling an ASCII text as UTF-8 is
    MvdV> correct.

    MvdV> Some old software may not recognise the kludge and so just assume the
    MvdV> default Fidonet character encoding. Which is ASCII, so all is well.
    MvdV> ;-)

    Your message suddenly brought to light a huge problem. And the problem isnʼt with the reader, but with the tosser. The thing is, the original message with a cyrillic UTF-8 name in the To: field, for some reason, didn’t make it into my JAM database for this echo.

    However, it passed through my station and reached 2:5075/21 without any issues. Thereʼs clearly a problem with the code of the hpt tosser from the Husky project: Iʼve already run into this once before—it has issues with cyrillic names in message headers, even in CP866. All in all, this will require further study of the issue.

    This shows that the technical readiness to use non-ASCII names in the “From” and “To” fields is still limited and can cause problems even with fairly modern software.

    Eugene

    ... It's full of stars!
    --- GoldED+/BSD 1.1.5-b20260828 (NetBSD 11.0 Intel Core Haswell)
    * Origin: FireFox Station (2:5075/35)
  • From Eugene Subbotin@2:5075/35 to Michiel van der Vlist on Sun Aug 30 04:07:14 2026
    Hello Michiel!

    Saturday August 29 2026 21:26, you wrote to Dmitry Lipatnikov:

    MvdV> So... How about enticing your NC and RC to participate in the UTF-8
    MvdV> nodelist project?

    I don't mind. I think my Russian name was listed there at one point, back when the UTF-8 nodelist was still a piece rather than being compiled by ZC.

    And I think we could probably get RC interested. Besides, he's already been testing the UTF-8 version of GoldED in a neighboring echo :)

    = UTF8.FTN.MESSAGING (2:5075/35) ==============================================
    From : Alex Barinov 2:5020/5452 Thu 27 Aug 26 01:08
    To : All
    Subj : test2 ===============================================================================
    Приветствую Вас, All!

    Слава императору на японском языке звучит как «Тэнно хэйка, бандзай!»
    (天皇陛下万歳), что дословно переводится как «Десять тысяч лет Его Величеству
    императору». Как это пишется и читается Кана / Иероглифы: 天皇陛下万歳Ромадзи
    (латиница): Tennō heika banzai! Произношение: Тэнноо хэйка бандзайСокращенный
    вариантЧасто используется только само слово «Бандзай!» (万歳 — «десять тысяч
    лет»), которое в историческом контексте выражает радость, победу или
    прославление монарха.

    Суп из собаки (посинтхан (보신탕; 補身湯), или кэджангук (개장국)) — блюдо
    национальной кухни Кореи, основным компонентом которого является мясо
    собаки[1]. Утверждается, что этот суп увеличивает мужскую силу[2].

    Если бы Кай из сказки Ханса Кристиана Андерсена собирал слово «вечность» на
    армянском, он бы складывал из ледяных осколков слово հավերժություն
    (haveržutʻyun).


    А вот с арабским как-то не очень красиво получается:

    أَعُوذُ بِاللَّهِ مِنَ الشَّيْطَانِ الرَّجِيمِ).


    Алексей Баринов

    E-Mail: aleksey.v.barinov AT gmail.com Skype: huba-huba Telegram: @huba715 [Team Бородатые] [F09F87BAF09F87A6] ===============================================================================

    Eugene

    ... It's full of stars!
    --- GoldED+/BSD 1.1.5-b20260828 (NetBSD 11.0 Intel Core Haswell)
    * Origin: FireFox Station (2:5075/35)
  • From Eugene Subbotin@2:5075/35 to Michiel van der Vlist on Sun Aug 30 05:04:32 2026
    Hello Michiel!

    Sunday August 30 2026 03:38, I wrote to you:

    Your message suddenly brought to light a huge problem. And the problem isnʼt with the reader, but with the tosser. The thing is, the original message with a cyrillic UTF-8 name in the To: field, for some reason, didn’t make it into my JAM database for this echo.

    However, it passed through my station and reached 2:5075/21 without
    any issues. Thereʼs clearly a problem with the code of the hpt tosser from the Husky project: Iʼve already run into this once before—it has issues with cyrillic names in message headers, even in CP866. All in
    all, this will require further study of the issue.

    This shows that the technical readiness to use non-ASCII names in the “From” and “To” fields is still limited and can cause problems even
    with fairly modern software.

    Following up on that point, a few more thoughts came to mind.

    Not only did my tosser, for some reason, fail to store a record with the To-field “Евгений Субботин” in the JAM database — and the reasons for this are still unknown — but there may be even more problems overall.

    For example, the name “Александр Христофоров” in UTF-8 will take up 41 bytes, which won’t fit within the 36 bytes defined in FTS-0001. This means that five bytes will be truncated along the way, which will corrupt the UTF-8 text. With GoldED+, exporting truncation will occur correctly — by character rather than by byte — but it may work differently on other systems.

    So there are far more problems, and they require a technical solution and standardization first before we can use such fields from the UTF-8 nodelist in headers.

    For example, FSP-1030 proposed using the kludges ^AUCSFROM:, ^AUCSTO:, and ^AUCSSUBJ: to include UTF-8 header fields in messages, but this was not adopted as a standard.

    Eugene

    ... It's full of stars!
    --- GoldED+/BSD 1.1.5-b20260829 (NetBSD 11.0 Intel Core Haswell)
    * Origin: FireFox Station (2:5075/35)
  • From Michiel van der Vlist@2:280/5555.1 to àóúÑ¡¿⌐ æπíí«γ¿¡ on Sun Aug 30 09:21:06 2026
    Hello �������,

    On 30 Aug 26 05:04, you wrote to me:

    Not only did my tosser, for some reason, fail to store a record with
    the To-field "������� �㡡�⨭" in the JAM database - and the reasons
    for this are still unknown

    OK... The area rules say here that UTF-8 is the preferred encoding here but it allows for others. So I have used this new control Y override to make this CP866. How does that work with your name in the header in Cyrillic?

    Just to rule out that it is not just the non ASCII that is the problem but UTF-8 specifically.

    BTW, I had to manually enter your cyrillic name in the header. For some reason copy/paste does not work for the header....


    Cheers, Michiel

    --- GoldED+/W32-MINGW 1.1.5-b20260829
    * Origin: Klein Schnoord (2:280/5555.1)
  • From Nick Boel@1:154/10 to Eugene Subbotin on Sun Aug 30 08:02:00 2026
    Hey Eugene!

    On Sun, 30 Aug 2026 05:04:32 +0300, you wrote:

    Following up on that point, a few more thoughts came to mind.

    Not only did my tosser, for some reason, fail to store a record with the To-field “Евгений Субботин” in the JAM database — and the reasons for
    this are still unknown — but there may be even more problems overall.

    For the record, I'm using the same (latest) version of hpt for Linux, and it seems that message stored in my JAM database just fine. I can grep the name right out of the .jhr file. Unless I'm not understanding this correctly..

    Regards,
    Nick

    ... Sarcasm: because beating people up is illegal.
    --- GoldED+/LNX 1.1.5-b20260830
    * Origin: _thePharcyde telnet://bbs.pharcyde.org (Wisconsin) (1:154/10)
  • From Carlos Navarro@2:341/234.1 to Michiel van der Vlist on Sun Aug 30 21:53:57 2026
    30 Aug 2026 09:21, you wrote to ������� �㡡�⨭:

    BTW, I had to manually enter your cyrillic name in the header. For
    some reason copy/paste does not work for the header....

    Alt-P only works for the body. In the header you have to right-click on your mouse (or press Alt+Space -> Edit -> Paste)

    Carlos

    --- GoldED+/W64-MSVC 1.1.5-b20260830
    * Origin: cyberiada (2:341/234.1)
  • From Michiel van der Vlist@2:280/5555 to Carlos Navarro on Sun Aug 30 22:02:20 2026
    Hello Carlos,

    On 30 Aug 26 21:53, you wrote to me:

    BTW, I had to manually enter your cyrillic name in the header.
    For some reason copy/paste does not work for the header....

    Alt-P only works for the body. In the header you have to right-click
    on your mouse (or press Alt+Space -> Edit -> Paste)

    I know and IIRC that worked before.

    But... It works with ASCII text. Not with the Cyrillic name I tried to copy/paste.

    Odd...

    Can you copy/paste cyrrylics into the header?


    Cheers, Michiel

    --- GoldED+/W32-MINGW 1.1.5-b20260830
    * Origin: Nieuw Schnøørd (2:280/5555)
  • From Eugene Subbotin@2:5075/35 to Nick Boel on Mon Aug 31 05:50:26 2026
    Hello Nick!

    Sunday August 30 2026 08:02, you wrote to me:

    Not only did my tosser, for some reason, fail to store a record
    with the To-field “Евгений Субботин” in the JAM database — and the
    reasons for this are still unknown — but there may be even more
    problems overall.

    For the record, I'm using the same (latest) version of hpt for Linux,
    and it seems that message stored in my JAM database just fine. I can
    grep the name right out of the .jhr file. Unless I'm not understanding this correctly..


    This issue likely only affects NetBSD: the previous bug also only affected that OS and didn't affect Linux. BSD isn't the most popular OS, after all :)

    Eugene

    ... It's full of stars!
    --- GoldED+/BSD 1.1.5-b20260830 (NetBSD 11.0 Intel Core Haswell)
    * Origin: FireFox Station (2:5075/35)
  • From Carlos Navarro@2:341/234.1 to Michiel van der Vlist on Mon Aug 31 14:18:41 2026
    30 Aug 2026 22:02, you wrote to me:

    Alt-P only works for the body. In the header you have to
    right-click on your mouse (or press Alt+Space -> Edit -> Paste)

    I know and IIRC that worked before.

    But... It works with ASCII text. Not with the Cyrillic name I tried to copy/paste.

    Odd...

    Yes...

    Can you copy/paste cyrrylics into the header?

    Yes, works for me on Win10. I'm pasting "Евгений Субботин" at the end of this message's subject.

    Carlos

    --- GoldED+/W64-MSVC 1.1.5-b20260830
    * Origin: cyberiada (2:341/234.1)
  • From Michiel van der Vlist@2:280/5555 to Eugene Subbotin on Mon Aug 31 17:50:03 2026
    Hello Eugene,

    On 30 Aug 26 05:04, you wrote to me:

    Your message suddenly brought to light a huge problem. And the
    problem isnʼt with the reader, but with the tosser. The thing is,
    the original message with a cyrillic UTF-8 name in the To: field,
    for some reason, didn’t make it into my JAM database for this
    echo.

    I see that in your part of the world cyrillics is used in the subject field. That doesn't cause similar problems?

    However, it passed through my station and reached 2:5075/21
    without any issues. Thereʼs clearly a problem with the code of
    the hpt tosser from the Husky project: Iʼve already run into this
    once before—it has issues with cyrillic names in message headers,
    even in CP866. All in all, this will require further study of the
    issue.

    Fmail doesn't have that problem, which shows that it is a tosser problem, not a poblem inherent with JAM. The hpt tosser is still under development and it is open source so suerely this particular problem can be solved.

    Anyway, that there are still some problem is not an argument to not participate in the UTF-8 nodelist project. One has to start somewhere and even if some problems remain, the project has already proven its value.

    This shows that the technical readiness to use non-ASCII names in
    the “From” and “To” fields is still limited and can cause
    problems even with fairly modern software.

    Problems can be solved...

    Following up on that point, a few more thoughts came to mind.

    Not only did my tosser, for some reason, fail to store a record with
    the To-field “Евгений Субботин” in the JAM database — and the reasons
    for this are still unknown — but there may be even more problems overall.

    ....

    For example, the name “Александр Христофоров” in UTF-8 will take up 41
    bytes, which won’t fit within the 36 bytes defined in FTS-0001.

    Yes, that is a problem. Fidonet was born and developed in he US and unfortunately the founding fathers did no thave the vision to foresee that one byte character sets would not be enough for he future. The rest of he digital word has evolved and 99,9% of the internet can handle UTF-8 multi byte characters.

    1) My name easely fits into the 36 byte limit, but if I have to fit it into 18 characers I would have to shorten it to "Michiel van Vlist". Or even "Michiel Vlist". Not nice but workable. Maybe "Александр Христофоров" can shorten it to "Алекс Христофоров".

    2)

    For example, FSP-1030 proposed using the kludges ^AUCSFROM:, ^AUCSTO:,
    and ^AUCSSUBJ: to include UTF-8 header fields in messages, but this
    was not adopted as a standard.

    Maybe we can make into a standard by just using it. That is how most standards came into being.

    3) The 36 byte limit as documente in FTS-0001 is for he stored format for the massage. But THAT is just a local thing that can be changed without having any impact on the rest of Fidonet. In the packed message, the field length is variable. The field is null terminated. Maybe simply ignoring the 36 byte limit on he packet has little or no impact at all. We could try. Rumour has it that some already did and it just worked.

    This means that five bytes will be truncated along the way, which will corrupt the UTF-8 text. With GoldED+, exporting truncation will occur correctly — by character rather than by byte — but it may work differently on other systems.

    4) Maybe. Or... we could just go ahead and see where the ship runs ashore.

    So there are far more problems, and they require a technical solution
    and standardization first before we can use such fields from the UTF-8 nodelist in headers.

    5) Don't wait for the standards to come down from haven. MAKE standards by going ahead and see what evolves into "common practise".

    6) So let us get R50 to join the UTF-8 nodelist project.


    Cheers, Michiel

    --- GoldED+/W32-MINGW 1.1.5-b20260830
    * Origin: Nieuw Schnøørd (2:280/5555)
  • From Michiel van der Vlist@2:280/5555 to Eugene Subbotin on Mon Aug 31 22:11:58 2026
    Hello Eugene,

    31 Aug 26 17:50, I wrote to you:

    For example, the name “Александр Христофоров” in UTF-8 will take
    up 41 bytes, which won’t fit within the 36 bytes defined in
    FTS-0001.

    7) Use compression on the From and To fields. There must be a substantial redundancy in the UTF-8 representation of Cyrillic (or other non ASCII) names. So if the encoding of the message is UTF-8 and the name is longer than 35 bytes and it contains non-ASCII, compress it. In all likehood the vast majority of names used in Fidonet will then fit in 36 bytes. (35 with the trailing zero.) On reading back, when it looks like "garbage", decompress it and if the result is well formed UTF-8 display that.

    Use a compression method that never has a zero byte in the result.


    Cheers, Michiel

    --- GoldED+/W32-MINGW 1.1.5-b20260830
    * Origin: Nieuw Schnøørd (2:280/5555)