Hybrid Environment for Robust Analysis of Language

整理文档很辛苦,赏杯茶钱您下走!

免费阅读已结束,点击下载阅读编辑剩下 ...

阅读已结束,您可以下载文档离线阅读编辑

资源描述

HERALDHybridEnvironmentforRobustAnalysisofLanguageDataAfzalBallim,GiovanniCorayandVincenzoPallottaSwissFederalInstituteofTechnologyLausanne{Ballim,Coray,Pallotta}@di.epfl.chApril19,1999AbstractThisprojectaddressestheproblemofperformingstructuralandsemanticanalysisofdatawherethesyntacticandsemanticmodelsofthedomainareinadequate,androbustmethodsmustbeemployedtoperforma“bestapproximation”toacompleteanalysis.Thisproblemisparticularlypertinentinthedomainoftextanalysis.Theabilitytodealwithlargeamountsofpossiblyill-formedorunforeseentextisoneoftheprincipalobjectivesofcurrentresearchinNaturalLanguageProcessingbycomputer(NLP),anabilitywhichisparticularlynecessaryforadvancedinformationextractionandretrievalfromlargetextualcorpora.Theresultsofthisworkcan,however,beappliedinotherdomainswhereamixofpartialgrammaticalandsemanticmodelsexist,suchasinimageanalysis.TheprojectbuildsonpreviousFNSRSprojectsbytheproposers.Inparticularitintegratesdiscourseanalysismeth-odsisadirectcontinuationofFNSRSprojectROTAwhichaddressedtheproblemsofdevelopingrobustgrammaticalanalysisonnoisyorpartiallydescribeddata.Whiletheproposershavehadmuchsuccessinthislatterprojectonthedevelopmentofefficientrobusttechniquesforgrammar-basedstructuralanalysisofdata,thesetechniquesmustbesupplementedbysemanticanalysis,becausemanyanalysisproblemscannotberesolvedinanyotherway.Thisprojectproposestheinvestigationofsuchmethodsandtheirintegrationwithstructuralanalysisintoahybridarchitecture.KeywordsRobustsemanticanalysis,Intelligentinformationextraction,Discourseanalysis.1IntroductionThedomainoftextanalysishasbeenchosenforitsrichnessatboththestructuralandsemanticlevel,aswellasthewidenumberofdomainsuponwhichittouches.Therapidexpansionofinformationsystemsatagloballevel,whichhasengenderedthenecessityforlarge-scaleautomaticanalysisoftextualdata,makesthisanareawherefundamentalresearchcanbeofgreatbenefit.Informationretrieval,datawarehousing,andknowledgemanagementareallareaswhichcanimmediatelyprofitfromprogressinthisdomain.1.1StateoftheArtFromaverysuperficialobservationofthehumanlanguageunderstandingprocess,itappearsclearthatnodeepcompe-tenceoftheunderlyingstructureofthespokenlanguageisrequiredinordertobeabletoprocessacceptablydistortedutterances.Ontheotherhand,themoreexperiencedisthespeaker,themoreprobableisasuccessfulunderstandingofthatdistortedinput.Howcanthiskindoffault-tolerantbehaviorbereproducedinanartificialsystembymeansofcom-putationaltechniques?Severalanswershavebeenproposedtothisquestionandmanysystemsimplementedsofar,butnooneofthemiscapableofdealingwithrobustnessasawhole.Psycholinguistictheoriesarebasedonanidealizedconceptoflanguageperformanceand/orcompetence,evenwhensta-tisticalmethodsareintroducedtoexplainphenomenawhicharehardlyunderstandablebymeansofaformaltheory.AsremarkedbyTedBriscoeinsection3.7of[ZU96]:“Despiteoverthreedecadesofresearcheffort,nopracticaldomain-independentparserofunrestrictedtexthasbeendeveloped”.Evenifthisstatementdatesbackto1996,duringtheselast1twoyearsnorealimprovementshasbeenmadeinachievingfullrobustnessforanNLPsystem.Howeverseveralattemptshasbeencarriedoutinordertoapproximatearobustbehavior.Themostcommonapproachistoextendaclassicaltheoryoflanguageunderstanding,oftenonlyataspecificlevel(mor-phologic,syntactic,semanticorpragmatic),tryingtoembodyacertaindegreeofrobustness.Thiskindofapproachmayseemreasonablyadequatesinceitisoftenbasedonasolidbackground,butitsuffersfromtheproblemofbeingbiasedandconstrainedbycanonicalapproachestoNLP.ThreedecadesofresearchinNLPandcomputationallinguisticscannotbediscarded,however,butitwouldbeusefultochangeperspectiveandseeiftheproblemofrobustnesscanbetackledfromadifferentpointofview.Anaturalconsequenceofthislaststatementwouldbethatoneshouldstartfromthescratchandusepasttechnology“byneeds”andnotbecauseof“trends”.Goinginmoredetail,thetwomainreasonsoffailuresinfollowingtheaboveapproachare:1.Sincehumansarecapableofdealingwithacceptableill-formedtextwithoutanydeepcompetenceoftheunderly-ingstructureofthelanguage,itseemsthatproposedtheoriesandsystemsarenotabletoperformanapproximatematchingbetweeninputandpre-definedstructures(atwhateverlinguisticlevel).2.Humansareabletocombinedifferentlevelofunderstandinginorderachieveanacceptableorevenpartialunder-standingoftheinputtext.Thus,afailureatacertainlevelcanberecoveredbyanotherlevelorasuitablecombi-nationoflevels.SystemsdesignedfollowingclassicalapproachestorobustnessinNLPareoftenmonolithicandnotconceivedtobeintegratedinadistributedcomputationalenvironmentwithbehavioursuchasthatshownbyhumans.Inthelastdecadetherehasbeenaproliferationofstochasticandprobabilisticmethodsappliedtoparsingtechnology.Unfortunately,asG.Gazdarpointedoutin[G.G96]thereareessentiallyfourproblemsthatcannotbesolvedbysimplyextendingstandardparsingtechniquesinthesuchadirection:1.Statisticalmethods1arenotabletoextractusefulprobabilitiesfrommodestysizedcorpora.2.The“SparseDataProblem”:N-gram-typesystems2areunabletodealwithdiscontinuousdependencieswhichper-vadenaturallanguageatanylinguisticlevel.Itisnotpossibletogiveanu

1 / 21
下载文档,编辑使用

©2015-2020 m.777doc.com 三七文档.

备案号:鲁ICP备2024069028号-1 客服联系 QQ:2149211541

×
保存成功