北京大学现代汉语语料库基本加工规范

整理文档很辛苦,赏杯茶钱您下走!

免费阅读已结束,点击下载阅读编辑剩下 ...

阅读已结束,您可以下载文档离线阅读编辑

资源描述

1TP391TheBasicProcessingofContemporaryChineseCorpusatPekingUniversitySPECIFICATIONYUShi-wenDUANHui-mingZHUXue-fengBingSWEN(InstituteofComputationalLinguistics,PekingUniversity,Beijing,100871)Abstract:TheInstituteofComputationalLinguistics,PekingUniversityhascompletedthebasicprocessingofacontemporaryChinesecorpusthathas27millionChineseCharacters.Inadditiontowordsegmentationandpart-of-speechtagging,theprocessinginvolvesthetaggingofpropernouns(personnames,placenames,organizationnamesandsoon),morphemesubcategoriesandthespecialusagesofverbsandadjectives.Thesuccessofthislarge-scalelanguageengineeringisattributedtotheSPECIFICATION,whichhadbeenmadebeforehandandwasbeingperfectedwhileinuse.WeareherebymakinganintroductiontotheSPECIFICATIONthroughthispublication,thusinvitingthecommentsfromalltheexpertsandourcolleaguesfortheimprovementofit.Keywords:contemporaryChinese;corpus;wordsegmentation;part-of-speechtagging;specification69483003973G1998030507486398519381219571219371219681042345*·67891011121314151617181920212223

1 / 23
下载文档,编辑使用

©2015-2020 m.777doc.com 三七文档.

备案号:鲁ICP备2024069028号-1 客服联系 QQ:2149211541

×
保存成功