A compression-based algorithm for Chinese word segmentation

Teahan, W.J.; Wen, Yingying; McNab, Rodger J.; Witten, Ian H.

A compression-based algorithm for Chinese word segmentation

Authors

Files

uow-cs-wp-1999-13.pdf (1.47 MB)

Permanent Link

https://hdl.handle.net/10289/1042

Abstract

The Chinese language is written without using spaces or other word delimiters. Although a text may be thought of as a corresponding sequence of words, there is considerable ambiguity in the placement of boundaries. Interpreting a text as a sequence of words is beneficial for some information retrieval and storage tasks: for example, full-text search, word-based compression, and keyphrase extraction. We describe a scheme that infers appropriate positions for word boundaries using an adaptive language model that is standard in text compression. It is trained on a corpus of pre-segmented text, and when applied to new text, interpolates word boundaries so as to maximize the compression obtained. This simple and general method performs well with respect to specialized schemes for Chinese language segmentation.

Citation

Teahan, W.J., Wen, Y., McNab, R. & Witten, H. (1999). A compression-based algorithm for Chinese word segmentation. (Working paper 99/13). Hamilton, New Zealand: University of Waikato, Department of Computer Science.

Type

Working Paper

Series name

Computer Science Working Papers

Date

1999-09

Publisher

Computer Science, University of Waikato

A compression-based algorithm for Chinese word segmentation

Authors

Files

Permanent Link

Publisher link

Rights

Abstract

Citation

Type

Series name

Date

Publisher

Degree

Type of thesis

Supervisor