{"id":83055,"date":"2020-05-28T01:42:15","date_gmt":"2020-05-27T23:42:15","guid":{"rendered":"https:\/\/prohoster.info\/blog\/administrirovanie\/kak-linuxovskij-sort-sortiruet-stroki"},"modified":"2020-05-28T01:42:15","modified_gmt":"2020-05-27T23:42:15","slug":"kak-linuxovskij-sort-sortiruet-stroki","status":"publish","type":"post","link":"https:\/\/prohoster.info\/en\/blog\/administrirovanie\/kak-linuxovskij-sort-sortiruet-stroki","title":{"rendered":"How Linux's sort command sorts strings","gt_translate_keys":[{"key":"rendered","format":"text"}]},"content":{"rendered":"<h1 id=\"vvedenie\">Introduction<\/h1>\n<p><\/p>\n<p>It all started with a short script meant to combine information about addresses <em>e-mail<\/em> of employees, obtained from a mailing list of users, with the positions of employees obtained from the HR database. Both lists were exported to text files in Unicode encoding <em>UTF-8<\/em> and saved with Unix line endings.<\/p>\n<p><\/p>\n<p>Contents <em>mail.txt<\/em><\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">Ivanov Andrey;ia@example.com<\/code><\/pre>\n<p><\/p>\n<p>Contents <em>buhg.txt<\/em><\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">Ivanova Alla;painter\nYolkina Ella;crane operator\nIvanov Andrey;locksmith\nAbakanov Mikhail;painter<\/code><\/pre>\n<p><\/p>\n<p>To merge, the files were sorted using a Unix command <em>sort<\/em> and fed into a Unix program <em>join<\/em>, which unexpectedly terminated with an error: <\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">$&gt; sort buhg.txt &gt; buhg.srt\n$&gt; sort mail.txt &gt; mail.srt\n$&gt; join buhg.srt mail.srt &gt; result\njoin: buhg.srt:4: is not sorted: Ivanov Andrey;locksmith<\/code><\/pre>\n<p><\/p>\n<p>Viewing the sorting result visually showed that overall, the sorting was correct, but in the case of matching male and female surnames, females were placed before males:<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">$&gt; sort buhg.txt\nAbakanov Mikhail;painter\nYolkina Ella;crane operator\nIvanova Alla;painter\nIvanov Andrey;locksmith<\/code><\/pre>\n<p><\/p>\n<p>Looks like a glitch in Unicode sorting or a manifestation of feminism in the sorting algorithm. The former, of course, is more plausible.<\/p>\n<p><noindex><a rel=\"nofollow\" name=\"habracut\"><\/a><\/noindex><\/p>\n<p>Let's put that aside for now <em>join<\/em> and focus on <em>sort<\/em>. Let's try to solve the problem by trial and error. To start, we will change the locale from <em>en_US<\/em> to <em>ru_RU<\/em>. For sorting, it would be sufficient to set the environment variable <em>LC_COLLATE<\/em>, but we won't be petty:<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">$&gt; LANG=ru_RU.UTF-8 sort buhg.txt\nAbakanov Mikhail;painter\nYolkina Ella;crane operator\nIvanova Alla;painter\nIvanov Andrey;locksmith<\/code><\/pre>\n<p><\/p>\n<p>Nothing changed.<\/p>\n<p><\/p>\n<p>Let's try to re-encode the files into a single-byte encoding: <\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">$&gt; iconv -f UTF-8 -t KOI8-R buhg.txt \n | LANG=ru_RU.KOI8-R sort \n | iconv -f KOI8-R -t UTF8<\/code><\/pre>\n<p><\/p>\n<p>Again, nothing changed.<\/p>\n<p><\/p>\n<p>Nothing can be done, we'll have to search for a solution on the internet. There\u2019s nothing specific about Russian surnames, but there are questions about other sorting oddities. Here\u2019s an example of such a problem: <noindex><a rel=\"nofollow\" href=\"https:\/\/serverfault.com\/questions\/95579\/unix-sort-treats-dash-characters-as-invisible\/95593\">Unix sort treats '-' (dash) characters as invisible<\/a><\/noindex>. In short, the strings \"a-b\", \"aa\", \"ac\" are sorted as \"aa\", \"a-b\", \"ac\".<\/p>\n<p><\/p>\n<p>The standard answer everywhere is: use the programmer's locale <em>\"C\"<\/em> and you will be happy. Let's try it:<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">$&gt; LANG=C sort buhg.txt\nYolkina Ella;crane operator\nAbakanov Mikhail;painter\nIvanov Andrey;locksmith\nIvanova Alla;lawyer<\/code><\/pre>\n<p><\/p>\n<p>Something has changed. The Ivanovs lined up in the correct order, but Yolkina slid somewhere. Let's return to the original task:<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">$&gt; LANG=C sort buhg.txt &gt; buhg.srt\n$&gt; LANG=C sort mail.txt &gt; mail.srt\n$&gt; LANG=C join buhg.srt mail.srt &gt; result<\/code><\/pre>\n<p><\/p>\n<p>It worked without errors, just as the internet promised. Despite the inclusion of Elkin in the first line.<\/p>\n<p><\/p>\n<p>The problem seems to be resolved, but just in case, let's try another Russian encoding \u2014 Windows encoding. <em>CP1251<\/em>:<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">$&gt; iconv -f UTF-8 -t CP1251 buhg.txt \n | LANG=ru_RU.CP1251 sort \n | iconv -f CP1251 -t UTF8 <\/code><\/pre>\n<p><\/p>\n<p>The sorting result, surprisingly, will match the locale. <em>\"C\"<\/em>, and the entire example, accordingly, passes without errors. It's some kind of mysticism.<\/p>\n<p><\/p>\n<p>I don't like mysticism in programming, as it usually hides errors. I will have to seriously address how it works <em>sort<\/em> and what it affects. <em>LC_COLLATE<\/em> .<\/p>\n<p><\/p>\n<p>In the end, I will try to answer questions:<\/p>\n<p><\/p>\n<ul>\n<li>why women's surnames were sorted incorrectly.<\/li>\n<li>why <em>LANG=ru_RU.CP1251<\/em> turned out to be equivalent. <em>LANG=C<\/em><\/li>\n<li>why there are <em>sort<\/em> and <em>join<\/em> different representations of the order of sorted lines.<\/li>\n<li>why all my examples have errors.<\/li>\n<li>finally, how to sort strings to my liking.<\/li>\n<\/ul>\n<p><\/p>\n<h1 id=\"sortirovka-v-yunikode\">Sorting in Unicode.<\/h1>\n<p><\/p>\n<p>The first stop will be technical report No. 10 titled <noindex><a rel=\"nofollow\" href=\"https:\/\/unicode.org\/reports\/tr10\/\">Unicode collation algorithm.<\/a><\/noindex> on the website <noindex><a rel=\"nofollow\" href=\"https:\/\/unicode.org\">unicode.org<\/a><\/noindex>. The report contains many technical details, so I'll allow myself to provide a summary of the main ideas.<\/p>\n<p><\/p>\n<p><em>Collation<\/em> \u2014 \"string comparison\" is the foundation of any sorting algorithm. The algorithms themselves can differ (\"bubble\", \"merge\", \"quick\"), but they all use pairwise string comparison to determine the order of their arrangement.<\/p>\n<p><\/p>\n<p>Sorting strings in natural language is quite a complex issue. Even in the simplest single-byte encodings, the order of letters in an alphabet that differs from the English Latin alphabet won\u2019t match the order of the numeric values that encode those letters. Thus, in the German alphabet, the letter <em>\u00d6<\/em> comes between <em>O<\/em> and <em>P<\/em>, while in the encoding <em>CP850<\/em> it falls between <em>\u00ff<\/em> and <em>\u00dc.<\/em>.<\/p>\n<p><\/p>\n<p>One can attempt to abstract from a specific encoding and consider \"ideal\" letters arranged in some order, as done in Unicode. Encodings <em>UTF8<\/em>, <em>UTF16<\/em> or single-byte <em>KOI8-R<\/em> (if a limited subset of Unicode is needed) will yield different numeric representations of letters, while still referring to the same elements of the base table. <\/p>\n<p><\/p>\n<p>It turns out that even by building a character table from scratch, we cannot assign a universal order to the characters. In various national alphabets that use the same letters, the order of those letters may differ. For example, in the French language, <em>\u00c6<\/em> will be considered a ligature and sorted as a string. <em>AE<\/em>In Norwegian, however, <em>\u00c6<\/em> will be a separate letter, which is placed after <em>Z<\/em>. By the way, besides ligatures like <em>\u00c6<\/em> there are letters represented by multiple symbols. For instance, in the Czech alphabet, there is the letter <em>Ch<\/em>, which stands between <em>H<\/em> and <em>I<\/em>.<\/p>\n<p><\/p>\n<p>In addition to differences in alphabets, there are also other national traditions that influence sorting. In particular, the question arises: in what order should words consisting of uppercase and lowercase letters follow in a dictionary? Additionally, punctuation marks can affect sorting. In Spanish, an inverted question mark is placed at the beginning of a question (<em>\u00bfTe gusta la m\u00fasica?<\/em>). In this case, it is obvious that questions should not be grouped into a separate cluster outside the alphabet, but how should strings with other punctuation marks be sorted?<\/p>\n<p><\/p>\n<p>I will not dwell on string sorting in languages that differ significantly from European ones. It is worth noting that in languages written from right to left or top to bottom, characters in strings are likely stored in reading order, and even in non-alphabetic scripts, there are ways to order strings character by character. For example, hieroglyphs can be ordered by stroke (<noindex><a rel=\"nofollow\" href=\"https:\/\/studychinese.ru\/kljuchi\/\">the keys of Chinese characters<\/a><\/noindex>) or by pronunciation. How emojis should be ordered, to be honest, I have no idea, but something can certainly be devised for them.<\/p>\n<p><\/p>\n<p>Based on the aforementioned features, the main requirements for comparing strings based on Unicode tables were formulated:<\/p>\n<p><\/p>\n<ul>\n<li>string comparison does not depend on the position of characters in the code table;<\/li>\n<li>sequences of characters that form a single character are brought to their canonical form (<em>A<\/em> + the upper circle is the same as <em>\u00c5<\/em>);<\/li>\n<li>); in string comparison, a character is considered in the context of the string and, if necessary, is combined with neighbors into a single comparison unit (<em>Ch<\/em> in Czech) or split into several (<em>\u00c6<\/em> in French);<\/li>\n<li>All national features (alphabet, uppercase\/lowercase letters, punctuation, order of writing styles) must be configurable down to manually assigning order (emoji);<\/li>\n<li>Comparison is important not only for sorting but also in many other places, for example, for defining the range of rows (substitution {A\u2026 z} in <em>bash<\/em>);<\/li>\n<li>comparison must be performed quickly enough.<\/li>\n<\/ul>\n<p><\/p>\n<p>Moreover, the authors of the report formulated comparison properties that algorithm developers should not rely on:<\/p>\n<p><\/p>\n<ul>\n<li>the comparison algorithm should not require a separate set of characters for each language (Russian and Ukrainian languages share most Cyrillic characters);<\/li>\n<li>comparison should not be based on the order of characters in Unicode tables;<\/li>\n<li>the weight of a string should not be an attribute of the string, as the same string in different cultural contexts can have different weights;<\/li>\n<li>the weights of strings can change upon merging or splitting (from <em>x<\/em> &lt; <em>y<\/em> it does not imply that <em>xz<\/em> &lt; <em>yz<\/em>);<\/li>\n<li>different strings with the same weights are considered equal from the sorting algorithm's perspective. Introducing additional ordering for such strings is possible, but it may degrade performance;<\/li>\n<li>in repeated sorts, strings with the same weights may swap places. Stability is a property of a specific sorting algorithm, not of the string comparison algorithm (see the previous point);<\/li>\n<li>sorting rules may change over time as cultural traditions are refined\/modified.<\/li>\n<\/ul>\n<p><\/p>\n<p>It is also stated that the comparison algorithm is unaware of the semantics of the processed strings. For instance, strings consisting only of digits should not be compared as numbers, and articles should not be removed from lists of English names (<em>Beatles, The<\/em>).<\/p>\n<p><\/p>\n<p>To meet all the specified requirements, a multi-level (essentially four-level) table sorting algorithm is proposed.<\/p>\n<p><\/p>\n<p>Firstly, characters in the string are converted to their canonical form and grouped into comparison units. Each comparison unit is assigned several weights corresponding to different levels of comparison. The weights of comparison units are elements of ordered sets (in this case, integers) that can be compared greater than-less than. A special value <em>IGNORED<\/em> (0x0) indicates that this unit does not participate in comparison at the corresponding level. String comparisons may be repeated several times, using the weights of the relevant levels. At each of these levels, the weights of the comparison units of the two strings are sequentially compared to each other.<\/p>\n<p><\/p>\n<p>In various implementations of the algorithm for different national traditions, the coefficient values may differ, but the Unicode standard includes a basic weight table \u2014 <em>\"Default Unicode Collation Element Table\"<\/em> (<em>DUCET<\/em>). I would like to note that setting the variable <em>LC_COLLATE<\/em> actually serves as an indication for choosing the weight table in the string comparison function.<\/p>\n<p><\/p>\n<p>Weight coefficients <em>DUCET<\/em> are structured as follows:<\/p>\n<p><\/p>\n<ul>\n<li>at the first level, all letters are brought to a single case, diacritical marks are discarded, and punctuation marks (not all) are ignored;<\/li>\n<li>at the second level, only diacritical marks are considered;<\/li>\n<li>at the third level, only case is taken into account;<\/li>\n<li>at the fourth level, only punctuation marks are taken into account.<\/li>\n<\/ul>\n<p><\/p>\n<p>Comparison occurs in several passes: first, the first level coefficients are compared; if the weights match, a reconsideration with the second level weights follows; then, possibly, the third and fourth levels.<\/p>\n<p><\/p>\n<p>Comparison ends when corresponding comparison units with different weights exist in the strings. Strings that have equal weights at all four levels are considered equal to each other.<\/p>\n<p><\/p>\n<p>This algorithm (with a lot of additional technical details) gave its name to report No. 10 \u2014 <em>\"Unicode Collation Algorithm\"<\/em> (<em>UCA<\/em>).<\/p>\n<p><\/p>\n<p>At this point, the sorting behavior from our example becomes a bit clearer. It would be good to compare it with the Unicode standard.<\/p>\n<p><\/p>\n<p>For testing implementations <em>UCA<\/em> there is a special <noindex><a rel=\"nofollow\" href=\"https:\/\/www.unicode.org\/Public\/UCA\/latest\/CollationTest.html\">test<\/a><\/noindex>, which uses <noindex><a rel=\"nofollow\" href=\"http:\/\/www.unicode.org\/Public\/UCA\/latest\/allkeys.txt\">a weight file<\/a><\/noindex>, implementing <em>DUCET<\/em>. In the weight file, you can find various curiosities. For example, it includes the order of Mahjong tiles and European dominoes, as well as the order of suits in a deck of cards (symbol <em>1F000<\/em> and beyond). The card suits are arranged according to bridge rules \u2014 \u2663\u2666\u2665\u2660, and the cards within each suit are in the sequence 10, 2, 3\u2026 K.<\/p>\n<p><\/p>\n<p>Manually verifying the correctness of string sorting according to <em>DUCET<\/em> would be quite tedious, but fortunately for us, there is an exemplary implementation library for working with Unicode \u2014 \"<noindex><a rel=\"nofollow\" href=\"http:\/\/site.icu-project.org\/\">International Components for Unicode<\/a><\/noindex>\" (<em>ICU<\/em>).<\/p>\n<p><\/p>\n<p>On the website of this library, developed in <em>IBM<\/em>, there are demo pages, including the <noindex><a rel=\"nofollow\" href=\"http:\/\/demo.icu-project.org\/icu-bin\/collation.html\">string comparison algorithm page<\/a><\/noindex>. We input our test strings with the default settings and, oh miracle, we get perfect Russian sorting.<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">Abakanov Mikhail; painter\nYolkina Ella; crane operator\nIvanov Andrey; plumber\nIvanova Alla; lawyer<\/code><\/pre>\n<p><\/p>\n<p>By the way, on the website <em>ICU<\/em> you can find clarifications about the algorithm's operation when processing punctuation marks. In the examples <noindex><a rel=\"nofollow\" href=\"http:\/\/userguide.icu-project.org\/collation\/faq\">Collation FAQ<\/a><\/noindex> apostrophes and hyphens are ignored.<\/p>\n<p><\/p>\n<p>Unicode helped us, but to find the reasons for the strange behavior <em>sort<\/em> downward API support (simultaneously with this in <em>Linux<\/em> we'll have to look elsewhere.<\/p>\n<p><\/p>\n<h1 id=\"sortirovka-v-glibc\">Sorting in glibc<\/h1>\n<p><\/p>\n<p>A quick look at the source codes of the utility <em>sort<\/em> from <em>GNU Core Utils<\/em> showed that the localization in the utility comes down to printing the current value of the variable <em>LC_COLLATE<\/em> when running in debug mode:<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">$ sort --debug buhg.txt &gt; buhg.srt\nsort: using \u2018en_US.UTF8\u2019 sorting rules<\/code><\/pre>\n<p><\/p>\n<p>String comparison is performed by the standard function <em>strcoll<\/em>, meaning all the interesting stuff is in the library <em>glibc<\/em>.<\/p>\n<p><\/p>\n<p>At <em>wiki<\/em> project <em>glibc<\/em> dedicated to string comparison <noindex><a rel=\"nofollow\" href=\"https:\/\/sourceware.org\/glibc\/wiki\/Locales#LC_COLLATE\">has one paragraph<\/a><\/noindex>. From this paragraph, we can understand that the <em>glibc<\/em> sorting is based on the algorithm we already know, <em>UCA<\/em> (<em>The Unicode collation algorithm<\/em>) and\/or on a closely related standard <em>ISO 14651<\/em> (<em>International string ordering and comparison<\/em>). Regarding the latter standard, it should be noted that on the website <noindex><a rel=\"nofollow\" href=\"https:\/\/standards.iso.org\/ittf\/PubliclyAvailableStandards\">standards.iso.org<\/a><\/noindex> <em>ISO 14651<\/em> it is officially declared public, but the corresponding link leads to a nonexistent page. Google returns several pages with links to official sites offering to purchase an electronic copy of the standard for a hundred euros, but on the third or fourth page of the search results, there are also direct links to <em>PDF<\/em>. Overall, the standard is practically indistinguishable from <em>UCA<\/em>, but it reads more boringly since it lacks vivid examples of national peculiarities of string sorting. <\/p>\n<p><\/p>\n<p>The most interesting information at <em>wiki<\/em> was a link to the <noindex><a rel=\"nofollow\" href=\"https:\/\/sourceware.org\/bugzilla\/show_bug.cgi?id=14095\">bug tracker<\/a><\/noindex> discussing the implementation of string comparison in <em>glibc<\/em>. From the discussion, one can learn that in <em>glibc<\/em> for string comparison a <em>ISO<\/em>template table is used <noindex><a rel=\"nofollow\" href=\"http:\/\/www.iso.org\/ittf\/ISO14651_2006_TABLE1_en.txt\">The Common Template Table<\/a><\/noindex> (<em>CTT<\/em>), whose address can be found in the appendix <em>A<\/em> of the standard <em>ISO 14651<\/em>. Between 2000 and 2015, this table in <em>glibc<\/em> did not have a maintainer and was quite different (at least externally) from the current version of the standard. From 2015 to 2018, there was an adaptation to the new version of the table, and at present, you have the chance to encounter both the new version of the table (<em>CentOS 8<\/em>), as well as the old version (<em>CentOS 7<\/em>). <\/p>\n<p><\/p>\n<p>Now that we have all the information about the algorithm and auxiliary tables, we can return to the original problem and understand how to properly sort strings in the Russian locale.<\/p>\n<p><\/p>\n<h1 id=\"iso-1465114652\">ISO 14651\/14652<\/h1>\n<p><\/p>\n<p>The source code of the table we are interested in <em>CTT<\/em> is located in the majority of distributions <em>Linux<\/em> in the directory <em>\/usr\/share\/i18n\/locales\/<\/em>. The table itself is located in the file <em>iso14651_t1_common<\/em>. Then this file is included in the file by the directive <em>copy iso14651_t1_common<\/em> which, in turn, is included in the national files, including in <em>. In most distributions<\/em>all source files are included in the basic installation, but if they are not available, you will have to install an additional package from the distribution. <em>en_US<\/em> and <em>ru_RU<\/em>may seem terribly verbose, with non-obvious rules for naming conventions, but if you break it down, it's quite simple. The structure is described in the standard <em>Linux<\/em> ISO 14652<\/p>\n<p><\/p>\n<p>File Structure <em>. In most distributions<\/em> , a copy of which can be downloaded from the website <em>open-std.org<\/em>. Another description of the file format can be read in <noindex><a rel=\"nofollow\" href=\"http:\/\/www.open-std.org\/JTC1\/SC22\/WG20\/docs\/n972-14652ft.pdf\">OpenGroup<\/a><\/noindex>. As an alternative to reading the standard, you can study the source texts of the function <noindex><a rel=\"nofollow\" href=\"https:\/\/pubs.opengroup.org\/onlinepubs\/9699919799\/basedefs\/V1_chap07.html\">the specifications<\/a><\/noindex> <em>POSIX<\/em> from <em>collate_read<\/em>glibc\/locale\/programs\/ld-collate.c <em>The file structure looks as follows:<\/em> downward API support (simultaneously with this in <em>By default, the symbol is used as the escape character, and the end of the line after the # symbol is a comment. Both characters can be overridden, which has been done in the new version of the table:<\/em>.<\/p>\n<p><\/p>\n<p>escape_char \/\ncomment_char %<\/p>\n<p><\/p>\n<p>In the file, there will be tokens in the format<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">escape_char \/ comment_char %<\/code><\/pre>\n<p><\/p>\n<p>The file will contain tokens in the format <em>\u2014 a hexadecimal digit). This is the hexadecimal representation of the Unicode code points in the<\/em> or <em>UCS-4<\/em> (where <em>x<\/em> UTF-32 <em>). All other elements in angle brackets (including<\/em> (<em>UTF-32<\/em>). All other elements in angle brackets (including <em>and similar), are considered plain string constants, having no special meaning outside of context.<\/em>, <em>indicates that the following contains data describing string comparisons.<\/em> First, the names for the weights in the comparison table and the names for symbol combinations are specified. Generally speaking, the two types of names belong to two different entities, but in the actual file, they are mixed. The names for weights are defined by the keyword<\/p>\n<p><\/p>\n<p>Line <em>LC_COLLATE<\/em> collating-symbol<\/p>\n<p><\/p>\n<p>First, names for the weights in the comparison table and names for combinations of characters are set. Generally speaking, two types of names belong to two different entities, but in the actual file, they are mixed. Weight names are specified with the keyword <em>collating-symbol<\/em> (comparison character), since when comparing Unicode characters with the same weights, they will be considered equivalent characters.<\/p>\n<p><\/p>\n<p>The total length of the section in the current revision of the file is about 900 lines. I pulled examples from several places to show the arbitrariness of names and some types of syntax.<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">LC_COLLATE\n\ncollating-symbol \ncollating-symbol \ncollating-symbol \ncollating-symbol \n...\ncollating-symbol \ncollating-symbol \ncollating-symbol \n...\ncollating-symbol ..\ncollating-symbol  % Guaranteed largest symbol value. Keep at end of this list\n...\ncollating-element  from \"\"\ncollating-element  from \"\"<\/code><\/pre>\n<p><\/p>\n<ul>\n<li><em>collating-symbol<\/em> registers a string <em>OSMANYA<\/em> in the weight names table <\/li>\n<li><em>collating-symbol ..<\/em> registers a sequence of names consisting of a prefix <em>S<\/em> and a hexadecimal numeric suffix from <em>1D000<\/em> up to <em>1D35F<\/em>.<\/li>\n<li><em>FFFF<\/em> downward API support (simultaneously with this in <em>collating-symbol<\/em> looks like a large unsigned integer in hexadecimal notation, but <em>&lt;SFFFF&gt;<\/em> is just a name that could look like <em>&lt;VERYBIGVAL&gt;<\/em> <\/li>\n<li>name <em>&lt;U0413&gt;<\/em> means a code point in the encoding <em>). All other elements in angle brackets (including<\/em><\/li>\n<li><em>collating-element  from \"\"<\/em> registers a new name for a pair of Unicode points. <\/li>\n<\/ul>\n<p><\/p>\n<p>When the weights' names are defined, the actual weights are assigned. Since only the greater-lesser relationships matter in comparison, the weights are determined by a simple enumeration sequence of names. \"Lighter\" weights are listed first, followed by \"heavier\" ones. I remind you that each Unicode character is assigned four different weights. Here they are compiled into a single ordered sequence. In theory, any symbolic name can be used at any of the four levels, but comments indicate that developers mentally categorize names by levels.<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">% Symbolic weight assignments\n\n% Third-level weight assignments\n\n\n\n\n...\n% Second-level weight assignments\n\n % COMBINING LOW LINE\n % COMBINING COMMA ABOVE\n % COMBINING REVERSED COMMA ABOVE\n...\n% First-level weight assignments\n % HORIZONTAL TABULATION \n % LINE FEED\n % VERTICAL TABULATION\n...\n % CYRILLIC SMALL LETTER DE\n % CYRILLIC SMALL LETTER KOMI DE\n % CYRILLIC SMALL LETTER DJE\n % CYRILLIC SMALL LETTER KOMI DJE\n % CYRILLIC SMALL LETTER GJE\n % CYRILLIC SMALL LETTER ZE WITH DESCENDER\n % CYRILLIC SMALL LETTER IE\n % CYRILLIC SMALL LETTER IE WITH BREVE\n % CYRILLIC SMALL LETTER UKRAINIAN IE\n % CYRILLIC SMALL LETTER ZHE<\/code><\/pre>\n<p><\/p>\n<p>Finally, the actual weight table.<\/p>\n<p><\/p>\n<p>The weight section is enclosed in rows with keywords. <em>order_start<\/em> and <em>order_end<\/em>. Additional parameters <em>order_start<\/em> determine the direction in which rows are viewed at each comparison level. By default, the parameter <em>forward<\/em>. The body of the section consists of rows that contain a character code and its four weights. The character code can be represented by the character itself, a code point, or a symbolic name defined earlier. Weights can also be specified by symbolic names, code points, or the characters themselves. If code points or symbols are used, their weight corresponds to the numeric value of the code point (the position in the Unicode table). Characters not explicitly specified (as I understand) are considered appended in the table with a primary weight that matches their position in the Unicode table. A special weight value <em>IGNORE<\/em> means that at the corresponding comparison level, this character is ignored.<\/p>\n<p><\/p>\n<p>To demonstrate the weight structure, I selected three quite obvious fragments:<\/p>\n<p><\/p>\n<ul>\n<li>characters that are completely ignored<\/li>\n<li>characters equivalent to the digit three at the first two levels<\/li>\n<li>the beginning of the Cyrillic alphabet, which does not contain diacritics and is therefore sorted mainly by the first and third levels.<\/li>\n<\/ul>\n<p><\/p>\n<pre><code class=\"plaintext\">order_start forward;forward;forward;forward,position\n IGNORE;IGNORE;IGNORE;IGNORE % NULL (in 6429)\n IGNORE;IGNORE;IGNORE;IGNORE % START OF HEADING (in 6429)\n IGNORE;IGNORE;IGNORE;IGNORE % START OF TEXT (in 6429)\n...\n ;;; % DIGIT THREE\n ;;; % FULLWIDTH DIGIT THREE\n ;;; % PARENTHESIZED DIGIT THREE\n ;;; % DIGIT THREE FULL STOP\n ;;<FONT>; % MATHEMATICAL BOLD DIGIT THREE\n...\n ;;; % CYRILLIC SMALL LETTER A\n ;;; % CYRILLIC CAPITAL LETTER A\n ;;; % CYRILLIC SMALL LETTER A WITH BREVE\n ;;; % CYRILLIC SMALL LETTER A WITH BREVE\n...\n ;;; % CYRILLIC SMALL LETTER BE\n ;;; % CYRILLIC CAPITAL LETTER BE\n ;;; % CYRILLIC SMALL LETTER VE\n ;;; % CYRILLIC CAPITAL LETTER VE\n...\norder_end<\/code><\/pre>\n<p><\/p>\n<p>Now we can return to sorting the examples from the beginning of the article. The catch lies in this part of the weight table:<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">IGNORE;IGNORE;IGNORE; % SPACE\n IGNORE;IGNORE;IGNORE; % EXCLAMATION MARK\n IGNORE;IGNORE;IGNORE; % QUOTATION MARK\n...<\/code><\/pre>\n<p><\/p>\n<p>It is clear that in this table, punctuation marks from the table <em>ASCII<\/em> (including spaces) are almost always ignored when comparing strings. The only exceptions are strings that match in every way except for punctuation marks that occur in matching positions. The strings from my example (after sorting) look like this for the comparison algorithm:<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">AbakanovMikhailPainter\nYolkinaEllacrane\nIvanovaAllaPainter\nIvanovAndreiJoiner<\/code><\/pre>\n<p><\/p>\n<p>Considering that in the weight table, uppercase letters in the Russian language come after lowercase letters (at the third level <em>&lt;CAP&gt;<\/em> heavier than <em>&lt;MIN&gt;<\/em>), the sorting appears absolutely correct.<\/p>\n<p><\/p>\n<p>When setting the variable <em>LC_COLLATE=C<\/em> a special table is loaded that defines byte-by-byte comparison<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">static const uint32_t collseqwc[] =\n{\n  8, 1, 8, 0x0, 0xff,\n  \\\/ * 1st-level table *\\\/ \n  6 * sizeof (uint32_t),\n  \\\/ * 2nd-level table *\\\/ \n  7 * sizeof (uint32_t),\n  \\\/ * 3rd-level table *\\\/ \n  L'x00', L'x01', L'x02', L'x03', L'x04', L'x05', L'x06', L'x07',\n  L'x08', L'x09', L'x0a', L'x0b', L'x0c', L'x0d', L'x0e', L'x0f',\n\n...\n  L'xf8', L'xf9', L'xfa', L'xfb', L'xfc', L'xfd', L'fe', L'xff'\n};<\/code><\/pre>\n<p><\/p>\n<p>Since in Unicode the code point for \u0401 is before A, the strings are sorted accordingly.<\/p>\n<p><\/p>\n<h1 id=\"tekstovye-i-dvoichnye-tablicy\">Text and binary tables<\/h1>\n<p><\/p>\n<p>It is obvious that string comparison is an extremely common operation, while parsing the table <em>CTT<\/em> is a rather expensive procedure. To optimize access to the table, it is compiled into binary form by the command <em>localedef<\/em>.<\/p>\n<p><\/p>\n<p>The command <em>localedef<\/em> takes as parameters a file with the table of national features (option <em>-i<\/em>), in which all characters are represented by Unicode points, and a file of correspondence between Unicode points and characters of a specific encoding (option <em>-f<\/em>). As a result of its execution, binary files for the locale are created, with a name specified in the last parameter.<\/p>\n<p><\/p>\n<p><em>Glibc<\/em> supports two binary file formats: \"traditional\" and \"modern\".<\/p>\n<p><\/p>\n<p>The traditional format implies that the locale name is the name of the subdirectory in <em>\/usr\/lib\/locale\/<\/em>. This subdirectory stores binary files <em>LC_COLLATE<\/em>, <em>LC_CTYPE<\/em>, <em>LC_TIME<\/em> and so on. The file <em>LC_IDENTIFICATION<\/em> contains the formal name of the locale (which may differ from the directory name) and comments.<\/p>\n<p><\/p>\n<p>The modern format assumes storing all locales in a single archive <em>\/usr\/lib\/locale\/locale-archive<\/em>, which is mapped into the virtual memory of all processes using it. <em>glibc<\/em>The locale name in the modern format undergoes some canonization\u2014only letters and digits are retained in the encoding names, converted to lowercase. Thus <em>ru_RU.KOI8-R<\/em>, will be preserved as <em>ru_RU.koi8r<\/em>.<\/p>\n<p><\/p>\n<p>Input files are searched in the current directory, as well as in the directories <em>\/usr\/share\/i18n\/locales\/<\/em> and <em>\/usr\/share\/i18n\/charmaps\/<\/em> for files <em>CTT<\/em> and encoding files respectively.<\/p>\n<p><\/p>\n<p>For example, the command<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">localedef -i ru_RU -f MAC-CYRILLIC ru_RU.MAC-CYRILLIC<\/code><\/pre>\n<p><\/p>\n<p>will compile the file <em>\/usr\/share\/i18n\/locales\/ru_RU<\/em> using the encoding file <em>\/usr\/share\/i18n\/charmaps\/MAC-CYRILLIC.gz<\/em> and save the result as <em>\/usr\/lib\/locale\/locale-archive<\/em> under the name <em>ru_RU.maccyrillic<\/em><\/p>\n<p><\/p>\n<p>If the variable <em>LANG=en_US.UTF-8<\/em> is set, then <em>glibc<\/em> it will look for binary locale files in the following sequence of files and directories:<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">\/usr\/lib\/locale\/locale-archive\n\/usr\/lib\/locale\/en_US.UTF-8\/\n\/usr\/lib\/locale\/en_US\/\n\/usr\/lib\/locale\/enUTF-8\/\n\/usr\/lib\/locale\/en\/<\/code><\/pre>\n<p><\/p>\n<p>If the locale is found in both traditional and modern formats, priority is given to the modern one.<\/p>\n<p><\/p>\n<p>You can view the list of compiled locales with the command <em>locale -a<\/em>.<\/p>\n<p><\/p>\n<h1 id=\"podgotovka-svoey-tablicy-sravneniya\">Preparing your own collation table<\/h1>\n<p><\/p>\n<p>Now, armed with knowledge, you can create your own perfect string collation table. This table should correctly compare Russian letters, including the letter \u0401, while considering punctuation marks according to the table <em>ASCII<\/em>.<\/p>\n<p><\/p>\n<p>The process of preparing your sorting table consists of two stages: editing the weight table and compiling it into binary form with the command <em>localedef<\/em>.<\/p>\n<p><\/p>\n<p>To adjust the comparison table with minimal editing effort, the format <em>open-std.org<\/em> provides for sections that adjust the weights of the existing table. The section starts with the keyword <em>reorder-after<\/em> and an indication of the position after which the replacement occurs. The section ends with the line <em>reorder-end.<\/em>If it's necessary to correct several areas of the table, a section is created for each of these areas.<\/p>\n<p><\/p>\n<p>I copied the new versions of the files <em>iso14651_t1_common<\/em> and <em>ru_RU<\/em> from the repository <em>glibc<\/em> to my home directory ~\/ .local \/ share \/ i18n \/ locales \/ and slightly edited the section <em>LC_COLLATE<\/em> downward API support (simultaneously with this in <em>ru_RU<\/em>. The new versions of the files are fully compatible with my version <em>glibc<\/em>. If you want to use the old versions of the files, you will have to change the symbolic names and the location from which the replacement starts in the table.<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">LC_COLLATE\n% Copy the template from ISO\/IEC 14651\ncopy \"iso14651_t1\"\nreorder-after \n ;; % SPACE\n ;; % EXCLAMATION MARK\n ;; % QUOTATION MARK\n...\n ;; % RIGHT CURLY BRACKET\n ;; % TILDE\nreorder-end\nEND LC_COLLATE<\/code><\/pre>\n<p><\/p>\n<p>In fact, the fields would need to be changed in <em>LC_IDENTIFICATION<\/em> so that they point to the locale <em>ru_MY<\/em>, but in my example this was not needed as I excluded locales from the archive search <em>locale-archive<\/em>.<\/p>\n<p><\/p>\n<p>To <em>localedef<\/em> worked with files in my folder through the variable <em>I18NPATH<\/em> you can add an additional directory for input file search, and the directory for saving binary files can be specified as a path with slashes:<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">$&gt; I18NPATH=~\/.local\/share\/i18n localedef -i ru_RU -f UTF-8 ~\/.local\/lib\/locale\/ru_MY.UTF-8<\/code><\/pre>\n<p><\/p>\n<p><em>POSIX<\/em> assumes that in <em>LANG<\/em> you can write absolute paths to directories with locale files starting with a forward slash, but <em>glibc<\/em> downward API support (simultaneously with this in <em>Linux<\/em> all paths are counted from the base directory, which can be overridden by the variable <em>LOCPATH<\/em>. After setting <em>LOCPATH=~\/.local\/lib\/locale\/<\/em> all files related to localization will only be searched in my folder. The locale archive when the variable is set <em>LOCPATH<\/em> is ignored.<\/p>\n<p><\/p>\n<p>Here\u2019s the decisive test:<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">$&gt; LANG=ru_MY.UTF-8 LOCPATH=~\/.local\/lib\/locale\/ sort buhg.txt\nAbakanov Mikhail;painter\nYolkina Ella;crane operator\nIvanov Andrey;plumber\nIvanova Alla;lawyer<\/code><\/pre>\n<p><\/p>\n<p>Hooray! We did it!<\/p>\n<p><\/p>\n<h1 id=\"rabota-nad-oshibkami\">Learning from Mistakes<\/h1>\n<p><\/p>\n<p>I already answered the questions about string sorting posed at the beginning, but there are still a couple of questions about errors \u2014 visible and invisible.<\/p>\n<p><\/p>\n<p>Let's go back to the original task.<\/p>\n<p><\/p>\n<p>And the program <em>sort<\/em> and the program <em>join<\/em> use the same string comparison functions from <em>glibc<\/em>. How did it happen that <em>join<\/em> gave a sorting error on the lines sorted by the command <em>sort<\/em> in the locale <em>en_US.UTF-8<\/em>? \u041e\u0442\u0432\u0435\u0442 \u043f\u0440\u043e\u0441\u0442: <em>sort<\/em> compares the entire string, while <em>join<\/em> only compares the key, which by default is the beginning of the string up to the first whitespace character. In my example, this led to an error message because the sorting of the first words in the lines did not match the sorting of the full strings.<\/p>\n<p><\/p>\n<p>The locale <em>\"C\"<\/em> ensures that in the sorted strings, the initial substrings up to the first space are also sorted, but this just masks the error. You can create such data (people with the same last names but different first names) that without an error message would yield incorrect results in file merging. If we want that <em>join<\/em> When merging file lines by full name, the correct approach is to explicitly specify the field separator and sort by the key field, not by the entire line. In this case, both the merge will proceed correctly, and there will be no errors in any locale:<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">$&gt; sort -t ; -k 1 buhg.txt &gt; buhg.srt\n$&gt; sort -t ; -k 1 mail.txt &gt; mail.srt\n$&gt; join -t ; buhg.srt mail.srt &gt; result<\/code><\/pre>\n<p><\/p>\n<p>Successfully completed example in encoding <em>CP1251<\/em> contains another error. The thing is, in all the distributions I know of <em>Linux<\/em> the packages lack a compiled locale <em>ru_RU.CP1251<\/em>. If a compiled locale is not found, then <em>sort<\/em> it silently uses byte-by-byte comparison, which is what we observed.<\/p>\n<p><\/p>\n<p>By the way, there is another small glitch related to the unavailability of compiled locales. The command <em>LOCPATH=\/tmp locale -a<\/em> will list all locales in <em>locale-archive<\/em>, but with the variable set <em>LOCPATH<\/em> for all programs (including the <em>locale<\/em>) those locales will be unavailable.<\/p>\n<p><\/p>\n<pre><code class=\"plaintext\">$&gt; LOCPATH=\/tmp locale -a | grep en_US\nlocale: Cannot set LC_CTYPE to default locale: No such file or directory\nlocale: Cannot set LC_MESSAGES to default locale: No such file or directory\nlocale: Cannot set LC_COLLATE to default locale: No such file or directory\nen_US\nen_US.iso88591\nen_US.iso885915\nen_US.utf8\n\n$&gt; LC_COLLATE=en_US.UTF-8 sort --debug\nsort: using \u2018en_US.UTF-8\u2019 sorting rules\n\n$&gt; LOCPATH=\/tmp LC_COLLATE=en_US.UTF-8 sort --debug\nsort: using simple byte comparison<\/code><\/pre>\n<p><\/p>\n<h1 id=\"zaklyuchenie\">Conclusion<\/h1>\n<p><\/p>\n<p>If you are a programmer who tends to think of strings as a collection of bytes, then your choice <em>LC_COLLATE=C<\/em>.<\/p>\n<p><\/p>\n<p>If you are a linguist or a dictionary compiler, it is better for you to compile your locale.<\/p>\n<p><\/p>\n<p>If you are a regular user, you just need to get used to the fact that the command <em>ls -a<\/em> outputs files that start with a dot, mixed with files that start with a letter, and <em>Midnight Commander<\/em>, which uses its internal functions to sort names, brings files starting with a dot to the beginning of the list.<\/p>\n<p><\/p>\n<h1 id=\"ssylki\">Links<\/h1>\n<p><\/p>\n<p><noindex><a rel=\"nofollow\" href=\"https:\/\/unicode.org\/reports\/tr10\/\">Report No. 10 Unicode collation algorithm <\/a><\/noindex><\/p>\n<p><\/p>\n<p><noindex><a rel=\"nofollow\" href=\"http:\/\/www.unicode.org\/Public\/UCA\/latest\/allkeys.txt\">Character weights on unicode.org <\/a><\/noindex><\/p>\n<p><\/p>\n<p><noindex><a rel=\"nofollow\" href=\"http:\/\/userguide.icu-project.org\/intro\"><em>ICU<\/em> \u2014 implementation of the library for working with Unicode from IBM. <\/a><\/noindex><\/p>\n<p><\/p>\n<p><noindex><a rel=\"nofollow\" href=\"http:\/\/demo.icu-project.org\/icu-bin\/collation.html\">Sorting test using <em>ICU<\/em> <\/a><\/noindex><\/p>\n<p><\/p>\n<p><noindex><a rel=\"nofollow\" href=\"http:\/\/www.iso.org\/ittf\/ISO14651_2006_TABLE1_en.txt\">Character weights in <em>ISO 14651<\/em> <\/a><\/noindex><\/p>\n<p><\/p>\n<p><noindex><a rel=\"nofollow\" href=\"http:\/\/www.open-std.org\/JTC1\/SC22\/WG20\/docs\/n972-14652ft.pdf\">Description of the file format with weights <em>open-std.org<\/em> <\/a><\/noindex><\/p>\n<p><\/p>\n<p><noindex><a rel=\"nofollow\" href=\"https:\/\/sourceware.org\/bugzilla\/show_bug.cgi?id=14095\">Discussion of string comparison in <em>glibc<\/em><\/a><\/noindex><\/p>\n<p>Source: <a content=\"nofollow\" rel=\"nofollow\" href=\"https:\/\/habr.com\/ru\/post\/503960\/\">habr.com<\/a> <\/p>","protected":false,"gt_translate_keys":[{"key":"rendered","format":"html"}]},"excerpt":{"rendered":"<p>\u0412\u0432\u0435\u0434\u0435\u043d\u0438\u0435 \u0412\u0441\u0451 \u043d\u0430\u0447\u0430\u043b\u043e\u0441\u044c \u0441 \u043a\u043e\u0440\u043e\u0442\u043a\u043e\u0433\u043e \u0441\u043a\u0440\u0438\u043f\u0442\u0430, \u043a\u043e\u0442\u043e\u0440\u044b\u0439 \u0434\u043e\u043b\u0436\u0435\u043d \u0431\u044b\u043b \u043e\u0431\u044a\u0435\u0434\u0438\u043d\u0438\u0442\u044c \u0438\u043d\u0444\u043e\u0440\u043c\u0430\u0446\u0438\u044e \u043e\u0431 \u0430\u0434\u0440\u0435\u0441\u0430\u0445 e-mail \u0441\u043e\u0442\u0440\u0443\u0434\u043d\u0438\u043a\u043e\u0432, \u043f\u043e\u043b\u0443\u0447\u0435\u043d\u043d\u044b\u0445 \u0438\u0437 \u0441\u043f\u0438\u0441\u043a\u0430 \u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u0442\u0435\u043b\u0435\u0439 \u043f\u043e\u0447\u0442\u043e\u0432\u043e\u0439 \u0440\u0430\u0441\u0441\u044b\u043b\u043a\u0438, \u0441 \u0434\u043e\u043b\u0436\u043d\u043e\u0441\u0442\u044f\u043c\u0438 \u0441\u043e\u0442\u0440\u0443\u0434\u043d\u0438\u043a\u043e\u0432, \u043f\u043e\u043b\u0443\u0447\u0435\u043d\u043d\u044b\u043c\u0438 \u0438\u0437 \u0431\u0430\u0437\u044b \u043e\u0442\u0434\u0435\u043b\u0430 \u043a\u0430\u0434\u0440\u043e\u0432. \u041e\u0431\u0430 \u0441\u043f\u0438\u0441\u043a\u0430 \u0431\u044b\u043b\u0438 \u044d\u043a\u0441\u043f\u043e\u0440\u0442\u0438\u0440\u043e\u0432\u0430\u043d\u044b \u0432 \u0442\u0435\u043a\u0441\u0442\u043e\u0432\u044b\u0435 \u0444\u0430\u0439\u043b\u044b \u0432 \u043a\u043e\u0434\u0438\u0440\u043e\u0432\u043a\u0435 \u042e\u043d\u0438\u043a\u043e\u0434 UTF-8 \u0438 \u0441\u043e\u0445\u0440\u0430\u043d\u0435\u043d\u044b \u0441 \u044e\u043d\u0438\u043a\u0441\u043e\u0432\u0441\u043a\u0438\u043c\u0438 \u043a\u043e\u043d\u0446\u0430\u043c\u0438 \u0441\u0442\u0440\u043e\u043a. \u0421\u043e\u0434\u0435\u0440\u0436\u0438\u043c\u043e\u0435 mail.txt \u0418\u0432\u0430\u043d\u043e\u0432 \u0410\u043d\u0434\u0440\u0435\u0439;ia@example.com \u0421\u043e\u0434\u0435\u0440\u0436\u0438\u043c\u043e\u0435 buhg.txt \u0418\u0432\u0430\u043d\u043e\u0432\u0430 \u0410\u043b\u043b\u0430;\u043c\u0430\u043b\u044f\u0440 \u0401\u043b\u043a\u0438\u043d\u0430 [&hellip;]<\/p>\n","protected":false,"gt_translate_keys":[{"key":"rendered","format":"html"}]},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[688],"tags":[],"class_list":["post-83055","post","type-post","status-publish","format-standard","hentry","category-administrirovanie"],"aioseo_notices":[],"aioseo_head":"\n\t\t<!-- All in One SEO 5.0.2 - aioseo.com -->\n\t<meta name=\"description\" content=\"\u0412\u0432\u0435\u0434\u0435\u043d\u0438\u0435 \u0412\u0441\u0451 \u043d\u0430\u0447\u0430\u043b\u043e\u0441\u044c \u0441 \u043a\u043e\u0440\u043e\u0442\u043a\u043e\u0433\u043e \u0441\u043a\u0440\u0438\u043f\u0442\u0430, \u043a\u043e\u0442\u043e\u0440\u044b\u0439 \u0434\u043e\u043b\u0436\u0435\u043d \u0431\u044b\u043b \u043e\u0431\u044a\u0435\u0434\u0438\u043d\u0438\u0442\u044c \u0438\u043d\u0444\u043e\u0440\u043c\u0430\u0446\u0438\u044e \u043e\u0431 \u0430\u0434\u0440\u0435\u0441\u0430\u0445 e-mail \u0441\u043e\u0442\u0440\u0443\u0434\u043d\u0438\u043a\u043e\u0432, \u043f\u043e\u043b\u0443\u0447\u0435\u043d\u043d\u044b\u0445 \u0438\u0437 \u0441\u043f\u0438\u0441\u043a\u0430 \u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u0442\u0435\u043b\u0435\u0439 \u043f\u043e\u0447\u0442\u043e\u0432\u043e\u0439 \u0440\u0430\u0441\u0441\u044b\u043b\u043a\u0438, \u0441.\" \/>\n\t<meta name=\"robots\" content=\"max-image-preview:large\" \/>\n\t<meta name=\"author\" content=\"Yuri Gagarin\"\/>\n\t<link rel=\"canonical\" href=\"https:\/\/prohoster.info\/en\/blog\/administrirovanie\/kak-linuxovskij-sort-sortiruet-stroki\" \/>\n\t<meta name=\"generator\" content=\"All in One SEO (AIOSEO) 5.0.2\" \/>\n\t\t<meta property=\"og:locale\" content=\"en_US\" \/>\n\t\t<meta property=\"og:site_name\" content=\"ProHoster | \u041a\u0443\u043f\u0438\u0442\u044c \u043d\u0430\u0434\u0435\u0436\u043d\u044b\u0439 \u0445\u043e\u0441\u0442\u0438\u043d\u0433 \u0434\u043b\u044f \u0441\u0430\u0439\u0442\u043e\u0432 \u0441 \u0437\u0430\u0449\u0438\u0442\u043e\u0439 \u043e\u0442 DDoS, VPS VDS \u0441\u0435\u0440\u0432\u0435\u0440\u044b\" \/>\n\t\t<meta property=\"og:type\" content=\"article\" \/>\n\t\t<meta property=\"og:title\" content=\"\ud83e\udd47\u041a\u0430\u043a Linux\u2019\u043e\u0432\u0441\u043a\u0438\u0439 sort \u0441\u043e\u0440\u0442\u0438\u0440\u0443\u0435\u0442 \u0441\u0442\u0440\u043e\u043a\u0438 | ProHoster\" \/>\n\t\t<meta property=\"og:description\" content=\"\u0412\u0432\u0435\u0434\u0435\u043d\u0438\u0435 \u0412\u0441\u0451 \u043d\u0430\u0447\u0430\u043b\u043e\u0441\u044c \u0441 \u043a\u043e\u0440\u043e\u0442\u043a\u043e\u0433\u043e \u0441\u043a\u0440\u0438\u043f\u0442\u0430, \u043a\u043e\u0442\u043e\u0440\u044b\u0439 \u0434\u043e\u043b\u0436\u0435\u043d \u0431\u044b\u043b \u043e\u0431\u044a\u0435\u0434\u0438\u043d\u0438\u0442\u044c \u0438\u043d\u0444\u043e\u0440\u043c\u0430\u0446\u0438\u044e \u043e\u0431 \u0430\u0434\u0440\u0435\u0441\u0430\u0445 e-mail \u0441\u043e\u0442\u0440\u0443\u0434\u043d\u0438\u043a\u043e\u0432, \u043f\u043e\u043b\u0443\u0447\u0435\u043d\u043d\u044b\u0445 \u0438\u0437 \u0441\u043f\u0438\u0441\u043a\u0430 \u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u0442\u0435\u043b\u0435\u0439 \u043f\u043e\u0447\u0442\u043e\u0432\u043e\u0439 \u0440\u0430\u0441\u0441\u044b\u043b\u043a\u0438, \u0441.\" \/>\n\t\t<meta property=\"og:url\" content=\"https:\/\/prohoster.info\/en\/blog\/administrirovanie\/kak-linuxovskij-sort-sortiruet-stroki\" \/>\n\t\t<meta property=\"og:image\" content=\"https:\/\/prohoster.info\/wp-content\/uploads\/2021\/11\/logo-350.jpg\" \/>\n\t\t<meta property=\"og:image:secure_url\" content=\"https:\/\/prohoster.info\/wp-content\/uploads\/2021\/11\/logo-350.jpg\" \/>\n\t\t<meta property=\"og:image:width\" content=\"350\" \/>\n\t\t<meta property=\"og:image:height\" content=\"350\" \/>\n\t\t<meta property=\"article:published_time\" content=\"2020-05-27T23:42:15+00:00\" \/>\n\t\t<meta property=\"article:modified_time\" content=\"2020-05-27T23:42:15+00:00\" \/>\n\t\t<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/prohoster\" \/>\n\t\t<meta property=\"article:author\" content=\"https:\/\/www.facebook.com\/prohoster\" \/>\n\t\t<!-- All in One SEO -->\n\n","aioseo_head_json":{"title":"\ud83e\udd47How Linux's sort sorts strings | ProHoster","description":"Introduction It all started with a short script that was supposed to merge email address information of employees, obtained from the mailing list users, with.","canonical_url":"https:\/\/prohoster.info\/en\/blog\/administrirovanie\/kak-linuxovskij-sort-sortiruet-stroki","robots":"max-image-preview:large","keywords":"","webmasterTools":{"miscellaneous":""},"schema":null,"og:locale":"en_US","og:site_name":"ProHoster | \u041a\u0443\u043f\u0438\u0442\u044c \u043d\u0430\u0434\u0435\u0436\u043d\u044b\u0439 \u0445\u043e\u0441\u0442\u0438\u043d\u0433 \u0434\u043b\u044f \u0441\u0430\u0439\u0442\u043e\u0432 \u0441 \u0437\u0430\u0449\u0438\u0442\u043e\u0439 \u043e\u0442 DDoS, VPS VDS \u0441\u0435\u0440\u0432\u0435\u0440\u044b","og:type":"article","og:title":"\ud83e\udd47\u041a\u0430\u043a Linux\u2019\u043e\u0432\u0441\u043a\u0438\u0439 sort \u0441\u043e\u0440\u0442\u0438\u0440\u0443\u0435\u0442 \u0441\u0442\u0440\u043e\u043a\u0438 | ProHoster","og:description":"\u0412\u0432\u0435\u0434\u0435\u043d\u0438\u0435 \u0412\u0441\u0451 \u043d\u0430\u0447\u0430\u043b\u043e\u0441\u044c \u0441 \u043a\u043e\u0440\u043e\u0442\u043a\u043e\u0433\u043e \u0441\u043a\u0440\u0438\u043f\u0442\u0430, \u043a\u043e\u0442\u043e\u0440\u044b\u0439 \u0434\u043e\u043b\u0436\u0435\u043d \u0431\u044b\u043b \u043e\u0431\u044a\u0435\u0434\u0438\u043d\u0438\u0442\u044c \u0438\u043d\u0444\u043e\u0440\u043c\u0430\u0446\u0438\u044e \u043e\u0431 \u0430\u0434\u0440\u0435\u0441\u0430\u0445 e-mail \u0441\u043e\u0442\u0440\u0443\u0434\u043d\u0438\u043a\u043e\u0432, \u043f\u043e\u043b\u0443\u0447\u0435\u043d\u043d\u044b\u0445 \u0438\u0437 \u0441\u043f\u0438\u0441\u043a\u0430 \u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u0442\u0435\u043b\u0435\u0439 \u043f\u043e\u0447\u0442\u043e\u0432\u043e\u0439 \u0440\u0430\u0441\u0441\u044b\u043b\u043a\u0438, \u0441.","og:url":"https:\/\/prohoster.info\/en\/blog\/administrirovanie\/kak-linuxovskij-sort-sortiruet-stroki","og:image":"https:\/\/prohoster.info\/wp-content\/uploads\/2021\/11\/logo-350.jpg","og:image:secure_url":"https:\/\/prohoster.info\/wp-content\/uploads\/2021\/11\/logo-350.jpg","og:image:width":350,"og:image:height":350,"article:published_time":"2020-05-27T23:42:15+00:00","article:modified_time":"2020-05-27T23:42:15+00:00","article:publisher":"https:\/\/www.facebook.com\/prohoster","article:author":"https:\/\/www.facebook.com\/prohoster"},"aioseo_meta_data":{"post_id":"83055","title":null,"description":null,"keywords":null,"keyphrases":null,"primary_term":null,"canonical_url":null,"og_title":null,"og_description":null,"og_object_type":"default","og_image_type":"default","og_image_url":null,"og_image_width":null,"og_image_height":null,"og_image_custom_url":null,"og_image_custom_fields":null,"og_video":null,"og_custom_url":null,"og_article_section":null,"og_article_tags":null,"twitter_use_og":false,"twitter_card":"default","twitter_image_type":"default","twitter_image_url":null,"twitter_image_custom_url":null,"twitter_image_custom_fields":null,"twitter_title":null,"twitter_description":null,"schema":{"blockGraphs":[],"customGraphs":[],"default":{"data":{"Article":[],"Course":[],"Dataset":[],"FAQPage":[],"Movie":[],"Person":[],"Product":[],"ProductReview":[],"Car":[],"Recipe":[],"Service":[],"SoftwareApplication":[],"WebPage":[]},"graphName":"","isEnabled":true},"graphs":[]},"schema_type":null,"schema_type_options":null,"pillar_content":false,"robots_default":true,"robots_noindex":false,"robots_noarchive":false,"robots_nosnippet":false,"robots_nofollow":false,"robots_noimageindex":false,"robots_noodp":false,"robots_notranslate":false,"robots_max_snippet":null,"robots_max_videopreview":null,"robots_max_imagepreview":"large","priority":null,"frequency":null,"local_seo":null,"seo_analyzer_scan_date":null,"breadcrumb_settings":null,"limit_modified_date":false,"reviewed_by":null,"ai":null,"created":"2021-02-28 15:26:22","updated":"2022-09-27 19:16:08","focus_keyword":null,"additional_keywords":null,"truseo_locale":null},"gt_translate_keys":[{"key":"link","format":"url"}],"_links":{"self":[{"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/posts\/83055","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/comments?post=83055"}],"version-history":[{"count":0,"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/posts\/83055\/revisions"}],"wp:attachment":[{"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/media?parent=83055"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/categories?post=83055"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/tags?post=83055"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}