Template talk:Unichar
Add topic
| This template does not require a rating on Wikipedia's content assessment scale. It is of interest to the following WikiProjects: | ||||||||
| ||||||||
Proposal: use Template:Char
[edit]Would it be good to place the character itself in {{char}}? jlwoodwa (talk) 06:43, 9 July 2023 (UTC)
- Although generally keen on char, I'd need to be convinced in this case. Char is used to "isolate" a glyph under discussion from the associated running text. In the output of unichar, that is usually clear.
- The only argument in favour that I can see is that, at present, unichar identifies the glyph by increasing its size and maybe the faint box used by char would be better? But conversely magnification makes it easier to "read".
- Did you have a particular case that provoked the proposal? 𝕁𝕄𝔽 (talk) 07:51, 9 July 2023 (UTC)
- It's clear to anyone who's familiar with the format, but I'm not sure it's as clear to a general reader, especially one who doesn't know what the "U+ stuff" means. I haven't noticed any specific problems that this would solve, I just think it's good to have a consistent format for "inline character literals" on Wikipedia. jlwoodwa (talk) 08:19, 9 July 2023 (UTC)
- So how would we handle this example: U+20E0 ⃠ COMBINING ENCLOSING CIRCLE BACKSLASH (which is already not handled terribly well). Likewise, Asiatic scripts present issues that don't occur to those of us only familiar with alphabetic scripts. A lot of development work has gone into this template to deal with these issues so changing it would not be trivial, given the need to verify many many test cases and rewrite to resolve anomalies. Annoyingly, one of the recent main developers, user:DePiep, is no longer available to advise. --𝕁𝕄𝔽 (talk) 10:20, 9 July 2023 (UTC)
- ⃠ seems to work just fine. I understand the difficulty of modifying such a convoluted and widely-used template, though. Since it sounds like it's not obviously a bad idea, I'll try the "obvious implementation" in the sandbox, and give an update here when it's working. jlwoodwa (talk) 10:35, 9 July 2023 (UTC)
- on Chrome, the symbol overruns the box (or the box underruns)... 𝕁𝕄𝔽 (talk) 13:43, 9 July 2023 (UTC)
- ... but then again it overruns the last digit of the codepoint right now. --𝕁𝕄𝔽 (talk) 13:45, 9 July 2023 (UTC)
- on Chrome, the symbol overruns the box (or the box underruns)... 𝕁𝕄𝔽 (talk) 13:43, 9 July 2023 (UTC)
- ⃠ seems to work just fine. I understand the difficulty of modifying such a convoluted and widely-used template, though. Since it sounds like it's not obviously a bad idea, I'll try the "obvious implementation" in the sandbox, and give an update here when it's working. jlwoodwa (talk) 10:35, 9 July 2023 (UTC)
- So how would we handle this example: U+20E0 ⃠ COMBINING ENCLOSING CIRCLE BACKSLASH (which is already not handled terribly well). Likewise, Asiatic scripts present issues that don't occur to those of us only familiar with alphabetic scripts. A lot of development work has gone into this template to deal with these issues so changing it would not be trivial, given the need to verify many many test cases and rewrite to resolve anomalies. Annoyingly, one of the recent main developers, user:DePiep, is no longer available to advise. --𝕁𝕄𝔽 (talk) 10:20, 9 July 2023 (UTC)
- It's clear to anyone who's familiar with the format, but I'm not sure it's as clear to a general reader, especially one who doesn't know what the "U+ stuff" means. I haven't noticed any specific problems that this would solve, I just think it's good to have a consistent format for "inline character literals" on Wikipedia. jlwoodwa (talk) 08:19, 9 July 2023 (UTC)
- I'm 2 years late, but what about allowing
|use=charso people can opt into this instead of changing the default behaviour? I don't think U+20E0 should be a blocker for this since it's already very broken in current unichar and in my opinion at least {{unichar/sandbox|20E0|use=char}} -> U+20E0 ⃠ COMBINING ENCLOSING CIRCLE BACKSLASH looks better than U+20E0 ⃠ COMBINING ENCLOSING CIRCLE BACKSLASH even if it is still broken. Warudo (talk) 21:30, 19 June 2025 (UTC)- (I forgot to mention that I added the feature to the sandbox for this discussion) Warudo (talk) 21:32, 19 June 2025 (UTC)
More flexibility in parameter 1
[edit]Occasionally, I'd like to use the unicode character itself as the parameter. For instance, for 🎴, I'd like {{unichar|🎴}} to produce U+1F3B4 🎴 FLOWER PLAYING CARDS. This would occasionally save me a short but slightly tedious round trip looking up the character code of a character I already have but I don't have the code of, and having the computer do this mapping for me seems quite doable using software (I don't know much about Wikipedia templates, though). Single characters 0-F/f can be exempt from this, of course, if their capacity to represent single-digit hexadecimal numbers from 0 to 15 is still important (although maybe it isn't, since most people write those like 000F or 0F anyway?).
While looking into this, I was reminded that the unichar template doesn't let you add the U+ prefix to the code in parameter 1. So, for instance, U+1F3B4 is an error. Apparently this is a common error for people to make, so maybe it should be detected and the U+ prefix should simply be stripped internally? Dingolover6969 (talk) 07:02, 19 October 2024 (UTC)
- I'm not sure how a reverse lookup like that could be easily accomplished in a Wikipedia template. It seems like something that ought to be possible since the computer obviously has this information, but I don't think you have access to the table that you'd need to do that. The best idea that comes to mind would be to generate a magic template list with a script or bot of some kind that hardcodes the table and then look it up from that. Andre🚐 07:08, 19 October 2024 (UTC)
- This is really easy to do with Lua modules. {{#invoke:ustring|codepoint|\🎴}} -> 127924 converts a unicode character to its corresponding code point. The problem is that adding support for this introduces ambiguity. Consider {{unichar|7}}. Should it return U+0007 <control-0007> or U+0037 7 DIGIT SEVEN? For this reason I oppose adding support for this feature. Instead, we should make a {{unichar2}} that accepts only unicode characters as parameters. Nickps (talk) 13:26, 21 February 2025 (UTC)
- I missed that Dingo already addressed this in the opening comment. I think that having 7 and 07 behave differently is unnecessarily confusing, so I've gone ahead and made {{unichar2}}. {{unichar2|🎴}} -> U+1F3B4 🎴 FLOWER PLAYING CARDS works as specified. Nickps (talk) 15:19, 21 February 2025 (UTC)
- Oh, nice work! Didn't know about that Lua module. Lua is awesome. There are definitely some templates that we did the old way that could be improved. Andre🚐 19:51, 21 February 2025 (UTC)
- Thanks! Yes, Lua is pretty useful, especially for stuff like this. From what I've seen, the two main reasons we don't use it more is because 1) not many people know Lua (I don't either) and 2) for some reason people really don't like calling modules from mainspace. Everything has to be a template.
- @Dingolover6969 Does this work for you? I'm pretty sure {{unichar2}} solves the problem you're having. Nickps (talk) 22:42, 21 February 2025 (UTC)
- Wow, very cool, thank you Nickps! I reckon I will find that template quite useful. Dingolover6969 (talk) 15:22, 22 February 2025 (UTC)
- You're welcome. I'm glad I could help. Nickps (talk) 22:07, 22 February 2025 (UTC)
- Wow, very cool, thank you Nickps! I reckon I will find that template quite useful. Dingolover6969 (talk) 15:22, 22 February 2025 (UTC)
- Oh, nice work! Didn't know about that Lua module. Lua is awesome. There are definitely some templates that we did the old way that could be improved. Andre🚐 19:51, 21 February 2025 (UTC)
- I missed that Dingo already addressed this in the opening comment. I think that having 7 and 07 behave differently is unnecessarily confusing, so I've gone ahead and made {{unichar2}}. {{unichar2|🎴}} -> U+1F3B4 🎴 FLOWER PLAYING CARDS works as specified. Nickps (talk) 15:19, 21 February 2025 (UTC)
- This is really easy to do with Lua modules. {{#invoke:ustring|codepoint|\🎴}} -> 127924 converts a unicode character to its corresponding code point. The problem is that adding support for this introduces ambiguity. Consider {{unichar|7}}. Should it return U+0007 <control-0007> or U+0037 7 DIGIT SEVEN? For this reason I oppose adding support for this feature. Instead, we should make a {{unichar2}} that accepts only unicode characters as parameters. Nickps (talk) 13:26, 21 February 2025 (UTC)
Can anyone see why combining characters have strange redirects?
[edit]- U+0302 ̂ COMBINING CIRCUMFLEX ACCENT
- U+0303 ̃ COMBINING TILDE
There are redirect articles (created just now) for Combining circumflex accent and Combining tilde. Clicking on the nlinked names above will not take you to either of those redirects. 0302 takes you to the top of Circumflex, 0303 takes you to Nasal vowel, which is not the only use of the diacritic. In each case, the top of the article shows "Redirected from <symbol>": normally I would choose this to find the redirecting article and correct it, but that doesn't seem to be possible. (There is actually a redirect article at ̃ (that's an orphaned combining tilde in there). So two questions:
- Why is the content of the
nlink=being ignored in favour of an article named for the character itself? (Compare with {{unichar|0023|nlink=pound sign}} still gets you U+0023 # NUMBER SIGN, no ifs not buts. [So this behaviour may be a relic of earlier error handling?]) - How can the erroneous redirects be corrected? (because {{unichar}} is not the only way to reach them.
Any ideas? (I assume that these two cases are not unique.) 𝕁𝕄𝔽 (talk) 23:01, 17 February 2025 (UTC)
- What did you put as the nlink content? As I edit this wikipedia page, it's telling me that the source code is
* {{unichar|0302|nlink=}} * {{unichar|0303|nlink=}}, which produces U+0302 ̂ COMBINING CIRCUMFLEX ACCENT U+0303 ̃ COMBINING TILDE, which link to what you describe (as expected based on my reading of the documentation). So, that's weird, if you put something else in (as I think your comment implies). I'll try including{{unichar|0302|nlink=Combining circumflex accent}} {{unichar|0303|nlink=Combining tilde}}in this comment to see if they work or if the same bug(?) affects them. U+0302 ̂ COMBINING CIRCUMFLEX ACCENT U+0303 ̃ COMBINING TILDE. - With regards to the erroneous redirects, I'm not sure what everything should redirect to, but https://en.wikipedia.org/w/index.php?title=%CC%83&redirect=no and https://en.wikipedia.org/w/index.php?title=%CC%84&redirect=no get me to the relevant redirect articles, which seem like they can be then edited normally. I got to these by going to another redirect page and then pasting the unicode character, which I copied from a website that puts it into one's clipboard, into the url — the "Redirected from <symbol>" note that appears on the other pages is also unclickable for me.
- I'm on firefox. The html for the "Redirected from <symbol>" note looks normal to me when I inspect-element it (
<a href="/w/index.php?title=%CC%83&redirect=no" class="mw-redirect" title="̃">̃</a>vs<a href="/w/index.php?title=~&redirect=no" class="mw-redirect" title="~">~</a>for a regular tilde (ok, it doesn't look normal, but the html seems to be correct)), so I imagine this is just about the browser choosing not to render links clickable when they only occur on combining diacritics. In fact, I've just verified this is the way it works on both firefox and chrome using the following html:<html><head></head><body>k<a href="https://example.com">̃</a></body></html>which isn't clickable. It isn't even blue in Chrome! So if we want this to change (which seems like a reasonable change to desire), I think we would have to file bug reports with the browsers themselves. Or do some hack to work around it in Wikipedia's software. - Dingolover6969 (talk) 10:00, 18 February 2025 (UTC)
- OK, these filled-nlink ones I tried seem to work for me. Dingolover6969 (talk) 10:01, 18 February 2025 (UTC)
- Thank you, that allowed me to fix the immediate problem with those two (and revealed many more, which I will work through). I didn't know about the
title=%XY%99hack. - Coming back to the more general point of
nlink=: "if no parameter is specified, then the redirect is to an article with the same name as the Unicode canonical name" – or at least that was I believed should happen: I was wrong. I see now that what actually happens is the link goes to the an article whose name is just one character long, the character itself. For 'ordinary' characters, that is not a problem because they are accessible but these combining ones are not. - (Normally, we only need to use the nlink if we want a specific section or if the Wikipedia article name and the Unicode canonical name differ.)
- Do we need to add anything to the documentation? 𝕁𝕄𝔽 (talk) 11:04, 18 February 2025 (UTC)
- Maybe if there is a cwith= then it should use the name of the character as the link.
- It would also be really nice if Unicode character properties were used to cause cwith= to happen for any combining character rather than caller having to do it! Spitzak (talk) 18:56, 18 February 2025 (UTC)
- Glad to hear it! :)
- The nlink documentation is certainly a little confusing, what with its mentions of "using its canonical name", by which it means that the canonical name is linked, to the unicode character. I was able to figure out the state of affairs by close examination of a quadruply-nested bullet point caveat; but it could certainly be made more evident.
- Dingolover6969 (talk) 06:46, 20 February 2025 (UTC)
- Thank you, that allowed me to fix the immediate problem with those two (and revealed many more, which I will work through). I didn't know about the
- OK, these filled-nlink ones I tried seem to work for me. Dingolover6969 (talk) 10:01, 18 February 2025 (UTC)
Template styles
[edit]Can we convert this to use TemplateStyles?
While code point names are an exception from MOS:SMALLCAPS, this template forces small-caps in a way that isn't overridable by user styles, which means users who struggle with all-caps text (hi! 👋🏼) can't use our user CSS to change the presentation.
If we convert to template styles, users can override for their login, where necessary. — OwenBlacker (he/him; Talk) 10:06, 21 February 2025 (UTC)
- Isn't the text in the database all-caps? It seems like this won't really help anybody who can't read all-caps text. Spitzak (talk) 17:01, 21 February 2025 (UTC)
- {{unichar/name}} converts the text to lowercase with
text-transform: lowercase;. It then stacksfont-variant: small-caps;on top of that so the now lowercase letters appear as small caps. Warudo (talk) 12:39, 5 September 2026 (UTC)
- {{unichar/name}} converts the text to lowercase with
Done @OwenBlacker, I hope you can excuse the delay. Warudo (talk) 12:48, 5 September 2026 (UTC)
- Thank you. Frustratingly, though, Spitzak was correct; the source data in commons:Data:Unicode data/names/000.tab (for example) in indeed all-caps. It seems like this might be a more complicated issue to resolve. We could use
text-transform: capitalizeto transform codepoint names, but that brings its own issues... — OwenBlacker (he/him; Talk). I support Wiki Workers United ✊🏼 18:39, 6 September 2026 (UTC)text-transform: capitalizewould not work. All that does is convert the first character of each word to uppercase. In this case, it is already uppercase. The only we can really do with all-caps text istext-transform: lowercase. Warudo (talk) 20:49, 6 September 2026 (UTC)
- Thank you. Frustratingly, though, Spitzak was correct; the source data in commons:Data:Unicode data/names/000.tab (for example) in indeed all-caps. It seems like this might be a more complicated issue to resolve. We could use
Jingtian (井)
[edit]Our example documentation for note= is U+4E95 井 CJK UNIFIED IDEOGRAPH-4E95 (Jingtian). However, I don't see why that should get a note. Based on my research (googling stuff), 井 (jǐng) is not a name for 井田制度 (jǐngtián zhìdù) (although 井田 (jǐngtián) is, apparently), even though 井田制度 is named after 井. I have added jǐngtián to the page for 井, so U+4E95 井 CJK UNIFIED IDEOGRAPH-4E95 should be fine... or maybe U+4E95 井 CJK UNIFIED IDEOGRAPH-4E95 (jǐng). Paging User:JMF, who added this, and so may wish to weigh in. This was also discussed on Talk:井#Wrong_target?, back when that page used to redirect to jingtian, I assume. Dingolover6969 (talk) 09:00, 20 April 2025 (UTC)
- @Dingolover6969, would you move this over to the talk page of the article where you saw it, please? because I don't remember it and it doesn't look like a "feature" of this template. 𝕁𝕄𝔽 (talk) 09:13, 20 April 2025 (UTC)
- and the reason I don't remember it is because I've been framed. I assume you mean this diff at number sign?
- Ah, I understand. I was thinking of this Template:Unichar/doc diff, but presumably the number sign diff is the progenitor of the example "in the wild".
- In that case, there's I don't think there's any remaining difficulty to discuss; it's just a single other Wikipedian getting confused and not knowing about the consensus that was reached about 井. I'll just remove that note from the pages on which it occurs.
- Out of curiosity, I also chased down the same verbiage on the Sharp_(music) page, and found this diff which ultimately is a correction of this diff; so, it's just some cruft that's been circulating Wikipedia since the early days.
- Anyway, sorry to bother you/thanks for your time 🙂 Dingolover6969 (talk) 02:27, 21 April 2025 (UTC)
Cyrillic example of use/use2, which doesn't seem to work?
[edit]I've just added an example from Japanese that demonstrates the value of use=lang.
But I can't see what the existing example is intended to demonstrate?
{{unichar|0485|cwith=◌|use=script|use2=Cyrs}}→ U+0485 ◌҅ COMBINING CYRILLIC DASIA PNEUMATA
since it doesn't actually render the CYRILLIC DASIA PNEUMATA diacritic with the place-holder character (◌). [btw, {{unichar|0485|cwith=◌}} produces U+0485 ◌҅ COMBINING CYRILLIC DASIA PNEUMATA, so that doesn't work either.]
What should happen with the Cyrillic example? 𝕁𝕄𝔽 (talk) 16:56, 21 April 2025 (UTC)
- I think again this shows there should be an argument that is "arbitrary wiki markup to print instead of the character". This would replace the lang, image, size, cwith, and a ton of other argument bloat, and also allow access to stuff that is currently impossible, such as the "cwith for two characters", or "remove the emoji formatting". Spitzak (talk) 17:52, 21 April 2025 (UTC)
text format
[edit]can someone add the ability to force display a character as text? — kwami (talk) 03:21, 11 May 2025 (UTC)
- @Kwamikagami: Can you explain what you mean?
- Taking your initial as an example, U+004B K LATIN CAPITAL LETTER K displays the character ⟨K⟩ as text.
- U+00A9 © COPYRIGHT SIGN displays the character ⟨©⟩ as text
- Or do you mean an emoji? like U+1F604 😄 SMILING FACE WITH OPEN MOUTH AND SMILING EYES displays a text description. I assume you don't mean :-D
- the parameter
image=does the opposite of what you ask, so I assume that this is not what you mean. (We should document the circumstances where this is justified. The only legitimate reason that I can think of is when there is a new codepoint but the glyph is not yet widely implemented in computer fonts. )
- More info please. --𝕁𝕄𝔽 (talk) 15:26, 11 May 2025 (UTC)
- sure,
- at 42355 Typhon, we said about a planetary symbol,
- the emoji variant of the character isn't appropriate here; we would want to specify it as the text variant, but i don't see how to apply u+FE0E to force it to display as text
- thanks — kwami (talk) 17:39, 11 May 2025 (UTC)
- Although the Unicode spec (Miscellaneous Symbols and Pictographs Range: 1F300–1F5FF) shows a two-tailed glyph such as you want, how U+1F300 is rendered is a type designer's choice. Google has chosen to represent it as whirlpool in their default computer font for Chrome. Other vendors may have taken a different approach.
- IFF you can find a computer font that renders it the way you want, you can wrap the {{unichar}} call in
span style="font-familyetc. Compare the use of U+0067 g LATIN SMALL LETTER G with U+0067 g LATIN SMALL LETTER G to chose a open-tail or closed-tail ⟨g⟩ (done using {{serif|{{unichar|0067}}}} and {{sans-serif|{{unichar|0067}}}}). BUT you can't assume that readers have that font – indeed you should assume that they don't have that font. I'm afraid you are stuck with usingfile:Typhon symbol (fixed width).svgif you are to be sure that what you see is what they get. "Unicode specifies the code point, not the glyph": the standard provides a general semantic meaning for a given code point; the images in the chart are suggestions or illustrations and no more. - Unicode does have a mechanism called 'Variation Sequences' (accessed using special Variation Selector code points) to specify alternative glyphs for certain characters. These are often used for things like different styles of CJK characters or mathematical symbols. However, there isn't a variation sequence defined for U+1F300 to specifically request the line drawing or "text" style. The effect of U+FE0E (and U+FE0F) relies entirely on whether the font in use has defined a specific rendering for the base character when followed by these variation selectors.𝕁𝕄𝔽 (talk) 22:51, 11 May 2025 (UTC) extended 22:55, 11 May 2025 (UTC)
- ah, my bad. i thought 1F300 was one of the characters that unicode defined as having both text and emoji glyphs. still, there are other characters that are so defined. that's deprecated practice now, but is permanently enshrined for the characters where it was implemented. shouldn't our unichar template support Unicode-defined glyphs? — kwami (talk) 00:00, 12 May 2025 (UTC)
- You had best do a new section to formally request that enhancement (though I don't know who is going to do it). But here's a thought: could
cwith=be used to prepend the U+FE0E? Do you know of any convenient test cases? 𝕁𝕄𝔽 (talk) 09:46, 12 May 2025 (UTC)- Being able to specify a variation selector is very much needed, mostly to turn off emoji variants. However I still recommend we do a really simple field which is "print this instead". This would cover the font selection, size selection, replacement images, multiple combining characters, variation selector, and all the other stuff that is piling up here as many confusing options and requests for options. Spitzak (talk) 12:41, 12 May 2025 (UTC)
- JMF, because 'cwith' is prepended, you have to put the desired character there, and then it displays the text character correctly but labels it U+FE0E 🌀︎ VARIATION SELECTOR-15 — kwami (talk) 19:34, 12 May 2025 (UTC)
- Then that's a bug (if I can be unfair to call it a bug, given that this is a new use for cwith). If I use
cwith=◌(together with combining tilde, for example), I don't get0303 dotted circle
, I get U+0303 ◌̃ COMBINING TILDE. Also, as already covered, 1F300 doesn't have a variant. Let me try a real example and see what happens. Real soon now... 𝕁𝕄𝔽 (talk) 22:29, 12 May 2025 (UTC)- i think the problem is that the vs has to come after the character. — kwami (talk) 22:39, 12 May 2025 (UTC)
- Then that's a bug (if I can be unfair to call it a bug, given that this is a new use for cwith). If I use
could cwith= be used to prepend the U+FE0E?
. Yes, it could but it doesn't matter because the variation selector must be placed after the character it modifies. Unichar does not provide a way to do this. Warudo (talk) 22:44, 12 May 2025 (UTC)- D'oh!!!! 𝕁𝕄𝔽 (talk) 22:47, 12 May 2025 (UTC)
- So, I think we should just allow editors to put whatever character they want after the character they display with a parameter that works just like cwith. I added this in the sandbox and called it
|suffix=. {{unichar/sandbox|01F604|suffix=︎}} results in the desired U+1F604 😄︎ SMILING FACE WITH OPEN MOUTH AND SMILING EYES. The reason I think this approach is better than adding a dedicated parameter for VS15 alone like @Kwamikagami did here is that there are uses for other variation selectors on Wikipedia (see slashed zero where VS1 is used) which means that in the future we might need to add even more parameters if we go down this path. Warudo (talk) 23:39, 12 May 2025 (UTC)- sure, that's straightforward. we could use vs15 as an example in the documentation as to what the new param might be used for. — kwami (talk) 01:59, 13 May 2025 (UTC)
- I would still prefer, rather than bloating this up with yet more arguments, to add a single "print exactly this for the character" argument. Possibly it could substitute the actual unicode code point for 'X' or something in this string if you are concerned that this makes it too easy to print something else. The font/image/size/prefix/suffix/etc is getting too extreme, and you still have not fixed the ability to show a double-width diacritic over two dotted circles. Spitzak (talk) 08:52, 13 May 2025 (UTC)
- if you put the dotted ring in both 'cwith' and 'suffix', that should work — kwami (talk) 09:07, 13 May 2025 (UTC)
- The suffix parameter allows: {{unichar/sandbox|0360|cwith=◌|suffix=◌}} -> U+0360 ◌͠◌ COMBINING DOUBLE TILDE. So it actually does give us
the ability to show a double-width diacritic over two dotted circles.
- On the other hand, I really don't like the idea of a "print exactly this for the character" parameter. It could easily lead to misleading output because of a mistake. Warudo (talk) 13:22, 13 May 2025 (UTC)
- sandbox version works for article 42355 Typhon — kwami (talk) 19:06, 13 May 2025 (UTC)
- Since using the sandbox in the encyclopedia is very risky as any editor might come along and break it, I ported the change to the actual template and replaced the call in 42355 Typhon. Since this change solves two problems with the template (the VS and the double-width diacritic) I feel justified in doing it. WP:BRD still applies if Spitzak or anyone else feels strongly about this. Warudo (talk) 19:48, 13 May 2025 (UTC)
- Suffix looks like it works. For some reason in Safari it sometimes (but not always) draws different-sized circles. Spitzak (talk) 21:56, 13 May 2025 (UTC)
- could that be your display font? circle+diacritic might force use of a different font than the default for the circle itself — kwami (talk) 23:08, 13 May 2025 (UTC)
- Sorry to spoil the party, but on my system [Chrome on ChromeOS], {{unichar|01F604|suffix=︎}} and {{unichar|01F604}} both give me exactly the same thing, an emoji smiley. So I guess Google's Roboto doesn't have the text-format grapheme? Can anyone suggest a font that does? 𝕁𝕄𝔽 (talk) 23:24, 13 May 2025 (UTC)
- i think the noto fonts do — kwami (talk) 23:53, 13 May 2025 (UTC)
- Sorry to spoil the party, but on my system [Chrome on ChromeOS], {{unichar|01F604|suffix=︎}} and {{unichar|01F604}} both give me exactly the same thing, an emoji smiley. So I guess Google's Roboto doesn't have the text-format grapheme? Can anyone suggest a font that does? 𝕁𝕄𝔽 (talk) 23:24, 13 May 2025 (UTC)
- could that be your display font? circle+diacritic might force use of a different font than the default for the circle itself — kwami (talk) 23:08, 13 May 2025 (UTC)
- Suffix looks like it works. For some reason in Safari it sometimes (but not always) draws different-sized circles. Spitzak (talk) 21:56, 13 May 2025 (UTC)
- Since using the sandbox in the encyclopedia is very risky as any editor might come along and break it, I ported the change to the actual template and replaced the call in 42355 Typhon. Since this change solves two problems with the template (the VS and the double-width diacritic) I feel justified in doing it. WP:BRD still applies if Spitzak or anyone else feels strongly about this. Warudo (talk) 19:48, 13 May 2025 (UTC)
- I agree with Warudo: "display whatever you like" seems like a good idea at first but it is too likely to be misused. Same logic that led us to stop requiring editors to supply a name and instead to fetch the canonical name from Wikidata. 𝕁𝕄𝔽 (talk) 23:13, 13 May 2025 (UTC)
- I might even have gone too far with suffix. Technically since a variation selector changes the appearance of the character, we should be telling people that it is there. I added a note to 42355 Typhon that the VS is there although that is probably moot since the notion that U+1F300 🌀 CYCLONE can be used for Typhon, with or without VS15, appears to be WP:OR. I posted a message on the article talk page about that. Warudo (talk) 00:18, 14 May 2025 (UTC)
- I added a note in the docs so hopefully any future usage of variation selectors is not hidden. Warudo (talk) 00:24, 14 May 2025 (UTC)
- I might even have gone too far with suffix. Technically since a variation selector changes the appearance of the character, we should be telling people that it is there. I added a note to 42355 Typhon that the VS is there although that is probably moot since the notion that U+1F300 🌀 CYCLONE can be used for Typhon, with or without VS15, appears to be WP:OR. I posted a message on the article talk page about that. Warudo (talk) 00:18, 14 May 2025 (UTC)
- sandbox version works for article 42355 Typhon — kwami (talk) 19:06, 13 May 2025 (UTC)
- The suffix parameter allows: {{unichar/sandbox|0360|cwith=◌|suffix=◌}} -> U+0360 ◌͠◌ COMBINING DOUBLE TILDE. So it actually does give us
- if you put the dotted ring in both 'cwith' and 'suffix', that should work — kwami (talk) 09:07, 13 May 2025 (UTC)
- You had best do a new section to formally request that enhancement (though I don't know who is going to do it). But here's a thought: could
- ah, my bad. i thought 1F300 was one of the characters that unicode defined as having both text and emoji glyphs. still, there are other characters that are so defined. that's deprecated practice now, but is permanently enshrined for the characters where it was implemented. shouldn't our unichar template support Unicode-defined glyphs? — kwami (talk) 00:00, 12 May 2025 (UTC)
Peculiar combining behaviour
[edit]Does anybody understand U+20E0 COMBINING ENCLOSING CIRCLE BACKSLASH? Because it is making unichar have a nervous breakdown,
- With no cwith= this is the result: U+20E0 ⃠ COMBINING ENCLOSING CIRCLE BACKSLASH.
- On my system, that displays an oversized "NO ..." sign that overlaps the concluding zero of 20E0, the space between, and the C of COMBINING.
- With cwith=◌, this is the result: U+20E0 ◌⃠ COMBINING ENCLOSING CIRCLE BACKSLASH
- On my system, that displays as <normal space><dotted circle><double-width space> between 20E0 and COMBINING.
- With cwith=P, this is the result: U+20E0 P⃠ COMBINING ENCLOSING CIRCLE BACKSLASH
- On my system, that displays as <normal space><capital P><tofu><normal space> between 20E0 and COMBINING.
Does not compute, Captain! 𝕁𝕄𝔽 (talk) 23:08, 13 May 2025 (UTC)
- with the last, i see 'no parking' -- a capital 'p' with an overstruck circle and slash
- both 2 and 3 look ok [if a bit crowded]; only 1 looks bad
- i'm using firefox on linux — kwami (talk) 20:23, 14 May 2025 (UTC)
- Ok, looks like another Roboto artefact. I'll stop worrying about it. Thanks. 𝕁𝕄𝔽 (talk) 21:37, 14 May 2025 (UTC)
Some character names are not found by the template
[edit]Currently in CJK Unified Ideographs Extension I § Background one finds U+2ED9D CJK UNIFIED IDEOGRAPH-2ED9D and U+2EDE0 CJK UNIFIED IDEOGRAPH-2EDE0. The reason the names of the characters are not shown is a bug in {{#invoke:Unicode data|lookup}}. I have already made an edit request to have the bug fixed, but I'm also leaving a message here since this talk page is more watched than the module's talk page. Warudo (talk) 13:45, 15 June 2025 (UTC)
I've seen that a Variation Selector is displayed as reserved code point on the documentation (e.g. <reserved-FE0E>). Who can fix it? If there is some APIs to look up Unicode data from Wikimedia servers directly and could be used in Lua modules, rather than writing dozens of submodules, that might be nice choice. -- Great Brightstar (talk) 15:34, 29 July 2025 (UTC)
- I've fixed this by adding the variation selectors to c:Data:Unicode_data/names/00F.tab. This is a hack however and should be reverted once the proper fix is added to Module:Unicode data. Warudo (talk) 16:18, 29 July 2025 (UTC)
An odd request: name only
[edit]I would like to fix a perennial problem with people 'correcting' the List of typographical symbols and punctuation marks. My plan is to fetch the canonical name from Commons, like Unichar does. I thought I found a way using the sub-template, unichar/name
- LATIN CAPITAL LETTER A
- ({{Unichar/name|na=LATIN CAPITAL LETTER A|hval=0041}})
but it requires the supplied name
-
- ({{Unichar/name|na=|hval=0041}})
though it does not require the codepoint value
- LATIN CAPITAL LETTER A
- ({{Unichar/name|na=LATIN CAPITAL LETTER A}})
I tried creating a template that is a copy of Unichar/name, except witb the na[me] code copied from Unichar. It didn't give way to brute force and ignorance, Is there a generous soul in the house? 𝕁𝕄𝔽 (talk) 19:14, 1 August 2025 (UTC)
- What you need is {{#invoke:Unicode data|lookup|name|41}} which returns LATIN CAPITAL LETTER A. This is what Unichar uses. If you don't want to call Lua directly from mainspace, just make a wrapper for it. Warudo (talk) 19:29, 1 August 2025 (UTC)
- Just to be clear, I'm not suggesting adding another bell or whistle to {{unichar}} but maybe another template?
- Anyway, TYVM, almost there.
- ´ || ACUTE ACCENT (freestanding) || Apostrophe, Grave, Circumflex || Combining acute accent
- I just need to reformat to reduce the size:
- ´ || ACUTE ACCENT (freestanding) || Apostrophe, Grave, Circumflex || || Combining acute accent
- and then wlink it:
- ´ || ACUTE ACCENT (freestanding) || Apostrophe, Grave, Circumflex || || Combining acute accent
- Ah... Linking is ignored
. 𝕁𝕄𝔽 (talk) 23:18, 1 August 2025 (UTC)
- You should do it the other way around. {{resize|[[{{#invoke:Unicode data|lookup|name|00B4}}]]}} gives ACUTE ACCENT. Warudo (talk) 23:26, 1 August 2025 (UTC)
- PS: If people object to the all caps, you can use parser functions to convert to lowercase. acute accent or Acute accent are both possible. Warudo (talk) 23:36, 1 August 2025 (UTC)
- Interesting, I didn't appreciate that effect.
- Of course that exposes another problem: none (or almost none) of the all-caps names exist as redirects, nor should they. I'm coming round to the idea that merging columns 1 and 2 and using standard unichar (or unichar2) is going to be a lot easier.
- U+00B4 ´ ACUTE ACCENT (freestanding) || Apostrophe, Grave, Circumflex || || Combining acute accent
- I'll float the idea at the article talk page. I have a feeling that the inability to sort separately by canonical name will be a show stopper. And the leading U+nnnn may not be too popular. 𝕁𝕄𝔽 (talk) 23:36, 1 August 2025 (UTC)
- @JMF Seeing the timestamp in your reply, I'm assuming you didn't see my comment about the parser functions. The all caps should not be a problem. Warudo (talk) 23:42, 1 August 2025 (UTC)
- True, we had an edit conflict. Adding the parser functions into the mix suggests rather emphatically that a new template would be needed. What do you think? Would it be worth the effort? 𝕁𝕄𝔽 (talk) 23:59, 1 August 2025 (UTC)
- I don't know. The template itself is only a single line of codeso there are no problems there. You can make it if you want. Whether it's worth the effort depends on how easy it would be to get consensus for the change. Warudo (talk) 00:17, 2 August 2025 (UTC)
<includeonly>{{ucfirst:{{lc:{{#invoke:Unicode data|lookup|name|{{{1}}}}}}}}}</includeonly>
- For the record, I think it's a good idea. I'd like to see it implemented. Warudo (talk) 00:22, 2 August 2025 (UTC)
- For the benefit of innocent bystanders, we should have closed this out to record the fact that this discussion led to the creation of template:ucname, which is being used in List of typographical symbols and punctuation marks. 𝕁𝕄𝔽 (talk) 16:23, 5 September 2026 (UTC)
- For the record, I think it's a good idea. I'd like to see it implemented. Warudo (talk) 00:22, 2 August 2025 (UTC)
- I don't know. The template itself is only a single line of code
- True, we had an edit conflict. Adding the parser functions into the mix suggests rather emphatically that a new template would be needed. What do you think? Would it be worth the effort? 𝕁𝕄𝔽 (talk) 23:59, 1 August 2025 (UTC)
- @JMF Seeing the timestamp in your reply, I'm assuming you didn't see my comment about the parser functions. The all caps should not be a problem. Warudo (talk) 23:42, 1 August 2025 (UTC)
- You should do it the other way around. {{resize|[[{{#invoke:Unicode data|lookup|name|00B4}}]]}} gives ACUTE ACCENT. Warudo (talk) 23:26, 1 August 2025 (UTC)
Italic style removal
[edit]There has to be a way to remove the italicized font-style from Arabic script. An example on the article Pe (Semitic letter)#Central Asian variant, looking like this:
U+06A7 ڧ ARABIC LETTER QAF WITH DOT ABOVE
Whereas it must look like this:
U+06A7 ڧ ARABIC LETTER QAF WITH DOT ABOVE
The reason is: italicized and stylized Arabic fonts don't ensure correct rendering, in fact, nearly all Arabic fonts lack and italic or oblique style!
Currently, there is no way of overriding the italic style in such cases! --Esperfulmo (talk) 15:20, 11 January 2026 (UTC)
- @Esperfulmo The issue should be fixed now. Using {{lang}} with Unichar will now only add italics if the character is in the Latin script (which was the intended behaviour). Because this was the result of a bug that has now been fixed, I don't see the need to add a manual override.
- If you are interested in the technical details: By default, lang automatically adds italics to anything that is written with the Latin script (because of MOS:NONENGITALIC) and because unichar uses HTML entities, lang incorrectly acted as if everything that unichar sends to it is in the Latin script. Warudo (talk) 16:38, 11 January 2026 (UTC)
- Thanks. Now the aforementioned example looks fine. --Esperfulmo (talk) 16:49, 11 January 2026 (UTC)
Should we support no-break space in |cwith= again?
[edit]Originally, using a blank |cwith= parameter set the combining character to U+00A0 NO-BREAK SPACE and the size to 150%. However, per this request by @JMF, I changed it to just use the dotted circle (and today I noticed that I forgot to remove the 150% thing so for the past year |cwith= and |cwith=◌ were different sizes for no reason).
But anyway, my question is, do we want to keep things as is with only ◌ supported or should we change the documentation to also allow |para= ? At a glance I don't see anything wrong with U+0485 ҅ COMBINING CYRILLIC DASIA PNEUMATA. In fact, U+20E0 ⃠ COMBINING ENCLOSING CIRCLE BACKSLASH looks better than U+20E0 ◌⃠ COMBINING ENCLOSING CIRCLE BACKSLASH on my PC because the former does not overlap with the name. Warudo (talk) 10:04, 7 September 2026 (UTC)
- Although, funnily enough with the 20E0 example, the dotted circle version looks more correct on my phone, so I guess there's no winning here. Warudo (talk) 10:07, 7 September 2026 (UTC)
- The very large majority of use cases for cwith are for combining diacritics, so U+25CC ◌ DOTTED CIRCLE is the obvious default (and to change it at this stage would generate an Everest of work) - and my opinion remains that a diacritic in isolation is unhelpful (and difficult for readers with impaired sight to identify even its existence so that they can magnify it). Equally, editors have the option to specify a different or specific 'background' character in their
cwith=<my choice>for cases like Combining circle. - It doesn't help that Chrome doesn't currently handle combining enclosing whatsits, so I can't even see what you think works. 𝕁𝕄𝔽 (talk) 10:43, 7 September 2026 (UTC)
- Editors have the choice to do it but our documentation tells them it's deprecated and that it might stop working at any moment. That's all I'm interested in changing. I don't want to change the default, I just think that the doc page should acknowledge
|cwith= as a valid alternative for characters like U+20E0. - As for Chrome not supporting it, I can't take a screenshot on my PC right now because I just left my house but if you can't test it on Firefox I'll send a screenshot later. Warudo (talk) 10:58, 7 September 2026 (UTC)
- As promised, here is a screenshot
Warudo (talk) 14:51, 7 September 2026 (UTC)
- Ah, ok, I should have guessed that you meant Chrome on Windows.
- Yes, your suggestion makes sense. If I remember correctly, the current wording was written to discourage use of unichar to create a false codepoiht for a character [let's say, 'q with horn', U+0071 ̛q LATIN SMALL LETTER Q) that doesn't exist as a pre-composed character. So feel free to revise it to suggest a legitimate use for it while continuing to deprecate silliness.
- I guess we would also need to make a note that says that, for the present at least, what you see may not be what they get. 𝕁𝕄𝔽 (talk) 16:20, 7 September 2026 (UTC) tweaked to demonstrate the fake 'q with horn'. Of course it is less of a problem now that we ignore the supplied name and only trust Wikidata. --𝕁𝕄𝔽 (talk) 16:50, 7 September 2026 (UTC)
- On second thought I won't add anything to the docs at this time. This is because as I said above U+20E0 still looks wrong on my phone even with . So there's really no point telling people to use a workaround that only affects 1 character and only on some OSes. If I come across another case where this is needed and works well enough I might add it then.
- Another problem I'm running into is that we can't just say "◌ and are the only allowed inputs" because that puts them on equal footing. We don't want editors to write U+0300 ̀ COMBINING GRAVE ACCENT because it looks too similar to the non-combining U+0060 ` GRAVE ACCENT and it goes against the convention of using ◌ for combining characters. We'd have to say "use ◌ and only use as a last resort if ◌ looks wrong" but this is too much WP:CREEP for a single problematic character. Warudo (talk) 17:58, 7 September 2026 (UTC)
- As promised, here is a screenshot
- Editors have the choice to do it but our documentation tells them it's deprecated and that it might stop working at any moment. That's all I'm interested in changing. I don't want to change the default, I just think that the doc page should acknowledge