Introduction to the special issue on the web as co

整理文档很辛苦,赏杯茶钱您下走!

免费阅读已结束,点击下载阅读编辑剩下 ...

阅读已结束,您可以下载文档离线阅读编辑

资源描述

c®2003AssociationforComputationalLinguisticsIntroductiontotheSpecialIssueontheWebasCorpusAdamKilgarriff¤GregoryGrefenstetteyLexicographyMasterClassLtd.andITRIClairvoyanceCorporationUniversityofBrightonTheWeb,teemingasitiswithlanguagedata,ofallmannerofvarietiesandlanguages,invastquantityandfreelyavailable,isafabulouslinguists’playground.ThisspecialissueofComputationalLinguisticsexploreswaysinwhichthisdreamisbeingexplored.1.IntroductionTheWebisimmense,free,andavailablebymouseclick.Itcontainshundredsofbillionsofwordsoftextandcanbeusedforallmanneroflanguageresearch.Thesimplestlanguageuseisspellchecking.Isitspeculaterorspeculator?Googlegives67fortheformer(usefullysuggestingthelattermighthavebeenintended)and82,000forthelatter.Questionanswered.LanguagescientistsandtechnologistsareincreasinglyturningtotheWebasasourceoflanguagedata,becauseitissobig,becauseitistheonlyavailablesourceforthetypeoflanguageinwhichtheyareinterested,orsimplybecauseitisfreeandinstantlyavailable.ThemodeofworkhasincreaseddramaticallyfromastandingstartsevenyearsagowiththeWebbeingusedasadatasourceinawiderangeofresearchactivities:Thepapersinthisspecialissueformasampleofthebestofit.Thisintroductiontotheissueaimstosurveytheactivitiesandexplorerecurringthemes.WerstconsiderwhethertheWebisindeedacorpus,thenpresentahistoryofthethemeinwhichweviewtheWebasadevelopmentoftheempiricistturnthathasbroughtcorporacenterstageinthecourseofthe1990s.WebrieysurveytherangeofWeb-basedNLPresearch,thenpresentestimatesofthesizeoftheWeb,forEnglishandforotherlanguages,andasimplemethodfortranslatingphrases.NextweopenthePandora’sboxofrepresentativeness(concludingthattheWebisnotrepresentativeofanythingotherthanitself,butthenneitherareothercorpora,andthatmoreworkneedstobedoneontexttypes).WethenintroducethearticlesinthespecialissueandconcludewithsomethoughtsonhowtheWebcouldbeputatthelinguist’sdisposalrathermoreusefullythancurrentsearchenginesallow.1.1IstheWebaCorpus?ToestablishwhethertheWebisacorpusweneedtondout,discover,ordecidewhatacorpusis.McEneryandWilson(1996,page21)sayInprinciple,anycollectionofmorethanonetextcanbecalledacorpus::::Buttheterm“corpus”whenusedinthecontextofmodernlinguisticstendsmostfrequentlytohavemorespecicconnotationsthanthissimpledenitionprovidesfor.Thesemaybeconsideredun-¤LewesRd,Brighton,BN24JG,UK.E-mail:Adam.Kilgarriff@itri.brighton.ac.ukySuite700,5001BaumBlvd,Pittsburgh,PA15213-1854.E-mail:grefen@clairvoyancecorp.com334ComputationalLinguisticsVolume29,Number3derfourmainheadings:samplingandrepresentativeness,nitesize,machine-readableform,astandardreference.Wewouldliketoreclaimthetermfromtheconnotations.Manyofthecollectionsoftextsthatpeopleuseandrefertoastheircorpus,inagivenlinguistic,literary,orlanguage-technologystudy,donott.AcorpuscomprisingthecompletepublishedworksofJaneAustenisnotasample,norisitrepresentativeofanythingelse.Closertohome,ManningandSch¨utze(1999,page120)observe:InStatisticalNLP,onecommonlyreceivesasacorpusacertainamountofdatafromacertaindomainofinterest,withouthavinganysayinhowitisconstructed.Insuchcases,havingmoretrainingdataisnormallymoreusefulthananyconcernsofbalance,andoneshouldsimplyuseallthetextthatisavailable.Wewishtoavoidasmugglingofvaluesintothecriterionforcorpus-hood.McEneryandWilson(followingothersbeforethem)mixthequestion“Whatisacorpus?”with“Whatisagoodcorpus(forcertainkindsoflinguisticstudy)?”muddyingthesimplequestion“Iscorpusxgoodfortasky?”withthesemanticquestion“Isxacorpusatall?”Thesemanticquestionthenbecomesadistraction,alltoolikelytoabsorbenergiesthatwouldotherwisebeaddressedtothepracticalone.Sothatthesemanticquestionmaybesetaside,thedenitionofcorpusshouldbebroad.Wedeneacorpussimplyas“acollectionoftexts.”Ifthatseemstoobroad,theonequalicationweallowrelatestothedomainsandcontextsinwhichthewordisusedratherthanitsdenotation:Acorpusisacollectionoftextswhenconsideredasanobjectoflanguageorliterarystudy.Theanswertothequestion“Isthewebacorpus?”isyes.2.HistoryForchemistryorbiology,thecomputerismerelyaplacetostoreandprocessinfor-mationgleanedabouttheobjectofstudy.Forlinguistics,theobjectofstudyitself(inoneofitstwoprimaryforms,theotherbeingacoustic)isfoundoncomputers.Textisaninformationobject,andacomputer’sharddiskisasvalidaplacetogoforitsrealizationastheprintedpageoranywhereelse.Theone-million-wordBrowncorpusopenedthechapteroncomputer-basedlan-guagestudyintheearly1960s.Notingthesingularneedsoflexicographyforbigdata,inthe1970sSinclairandAtkinsinauguratedtheCOBUILDproject,whichraisedthethresholdofviablecorpussizefromonemillionto,bytheearly1980s,eightmillionwords(Sinclair1987).Tenyearson,Atkinsagaintooktheleadwiththedevelop-ment(from1988)oftheBritishNationalCorpus(BNC)(Burnard1995),whichraisedhorizonstenfoldonceagain,withits100millionwordsandwasinadditionwidelyavailableatlowcostandcoveredawidespectrumofvarietiesofcontemporaryBritishEnglish.1AsinallmattersZipan,logarithmicgraphpaperisrequired.Wherecorpussizeisconcerned,thestepsofinterestare1,10,100,:::,not1,2,3,:::Corporacrashedintocomputationallinguisticsatthe1989ACLmeeti

1 / 15
下载文档,编辑使用

©2015-2020 m.777doc.com 三七文档.

备案号:鲁ICP备2024069028号-1 客服联系 QQ:2149211541

×
保存成功