cleaner code in dal
This commit is contained in:
@@ -0,0 +1,19 @@
|
||||
# Copyright (c) 2006-2012 James Tauber and contributors
|
||||
#
|
||||
# Permission is hereby granted, free of charge, to any person obtaining a copy
|
||||
# of this software and associated documentation files (the "Software"), to deal
|
||||
# in the Software without restriction, including without limitation the rights
|
||||
# to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
||||
# copies of the Software, and to permit persons to whom the Software is
|
||||
# furnished to do so, subject to the following conditions:
|
||||
#
|
||||
# The above copyright notice and this permission notice shall be included in
|
||||
# all copies or substantial portions of the Software.
|
||||
#
|
||||
# THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||
# IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
||||
# FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
||||
# AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
||||
# LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
||||
# OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
|
||||
# THE SOFTWARE.
|
||||
@@ -0,0 +1,43 @@
|
||||
# pyuca: Python Unicode Collation Algorithm implementation
|
||||
(http://jtauber.com/blog/2006/01/27/python_unicode_collation_algorithm/)
|
||||
|
||||
This is my preliminary attempt at a Python implementation of the
|
||||
[Unicode Collation Algorithm (UCA)](http://unicode.org/reports/tr10/).
|
||||
I originally posted it to my blog in 2006 but it seems to get enough
|
||||
usage it really belongs here (and in PyPI).
|
||||
|
||||
What do you use it for? In short, sorting non-English strings properly.
|
||||
|
||||
The core of the algorithm involves multi-level comparison. For example,
|
||||
``café`` comes before ``caff`` because at the primary level, the accent
|
||||
is ignored and the first word is treated as if it were ``cafe``.
|
||||
The secondary level (which considers accents) only applies then to words
|
||||
that are equivalent at the primary level.
|
||||
|
||||
The Unicode Collation Algorithm and pyuca also support contraction and
|
||||
expansion. **Contraction** is where multiple letters are treated as a
|
||||
single unit. In Spanish, ``ch`` is treated as a letter coming between
|
||||
``c`` and ``d`` so that, for example, words beginning ``ch`` should
|
||||
sort after all other words beginnings with ``c``. **Expansion** is where
|
||||
a single letter is treated as though it were multiple letters. In German,
|
||||
``ä`` is sorted as if it were ``ae``, i.e. after ``ad`` but before ``af``.
|
||||
|
||||
## Here is how to use the ``pyuca`` module:
|
||||
``
|
||||
git clone https://github.com/jtauber/pyuca.git
|
||||
cd pyuca
|
||||
pip install pyuca
|
||||
``
|
||||
|
||||
**Usage example:**
|
||||
``
|
||||
from pyuca import Collator
|
||||
c = Collator("allkeys.txt")
|
||||
|
||||
sorted_words = sorted(words, key=c.sort_key)
|
||||
``
|
||||
|
||||
``allkeys.txt`` (1 MB) is available at
|
||||
|
||||
http://www.unicode.org/Public/UCA/latest/allkeys.txt
|
||||
|
||||
Reference in New Issue
Block a user