Changes

Full-Text: Japanese (edit)

Revision as of 13:56, 2 July 2020

38 bytes added , 13:56, 2 July 2020

no edit summary

This article is linked from the [[Full-Text]] page. It gives some insight into the implementation of the full-text features for Japanese text corpora. The Japanese version is [~~http~~https://files.basex.org/etc/ja-ft.pdf also available as PDF].~~Thank you to [http://blog.infinite.jp~~ The lexer was contributed by Toshio HIRAI~~] for integrating the lexer in BaseX!~~.

=Introduction=

The lexical analysis of Japanese documents is performed by [~~http~~https://igo.~~sourceforge~~osdn.jp/ Igo]. Igo is a ''morphological analyser'',and some of the advantages and reasons for using Igo are:* compatible with the results of a prominent morphological analyzer "MeCab"* it can use the dictionary distributed by the Project MeCab* the morphological analyzer is implemented in Java and is relatively fast

* Compatible with the results of a prominent morphological analyzer "MeCab".* It can use the dictionary distributed by the Project MeCab.* The morphological analyzer is implemented in Java and is relatively fast. Japanese tokenization will be activated in BaseX if Igo is found in theclasspath. [~~http~~https://~~en.sourceforge~~osdn.jpnet/projects/igo/releases/ igo-0.4.3.jar]of Igo is currently included in all distributions of BaseX.

In addition to the library, one of the following dictionary files must either be unzipped into the current directory, or into the <code>etc</code> sub-directory of the project’s [[Configuration#Home Directory|Home Directory]]:

* IPA Dictionary: ~~http~~https://files.basex.org/etc/ipadic.zip* NAIST Dictionary: ~~http~~https://files.basex.org/etc/naistdic.zip

=Lexical Analysis=

=Token Processing=

"Fullwidth" and "Halfwidth" (which is defined by[~~http~~https://unicode.org/Public/UNIDATA/EastAsianWidth.txt East Asian Width Properties])are not distinguished (this is the so-called ZENKAKU/HANKAKU problem). For example, <code>ＸＭＬ</code> and <code>XML</code> will be treatedas the same word. If documents are ''hybrid'', i.e. written in multiple languages,this is also helpful for some other options of the XQuery Full Text Specification,such as the [~~http~~https://www.w3.org/TR/xpath-full-text-10/#ftcaseoption Case] or the[~~http~~https://www.w3.org/TR/xpath-full-text-10/#ftdiacriticsoption Diacritics] ~~Option~~option.

=Stemming=

is returned for the following two types of queries: