An Approach for Improving Execution Performance in

血小红
1 ℃
2020-01-28

整理文档很辛苦，赏杯茶钱您下走！

还剩 ... 页未读，继续阅读 >>

免费阅读已结束，点击下载阅读编辑剩下 ... 页

阅读已结束，您可以下载文档离线阅读编辑

资源描述

AnApproachforImprovingExecutionPerformanceinInferenceNetworkBasedInformationRetrievalEricW.BrownTechnicalReport94–73September1994DepartmentofComputerScienceUniversityofMassachusettsAmherst,MA01003USAAbstractTheinferencenetworkretrievalmodelprovidestheabilitytocombineavarietyofretrievalstrategiesexpressedinarichquerylanguage.Whilethispoweryieldsimpressiveretrievaleffectiveness,italsopresentsbarrierstotheincorporationoftraditionaloptimizationtechniquesintendedtoimprovetheexecutionefﬁciency,orspeed,ofretrieval.Theessenceoftheseoptimizationtechniquesistheabilitytoidentifytheindexinginformationthatwillbemostusefulforevaluatingaquery,combinedwiththeabilitytoaccessfromdiskjustthatinformation.Isurveyavarietyoftechniquesintendedtoidentifyusefulindexinginformationandproposeanewstrategyforincorporatingthesetechniquesintotheinferencenetworkretrievalmodel.Additionally,Idescribeanarchitecturebasedonapersistentobjectstorethatprovidessophisticatedmanagementoftheindexingdata,allowingjustthedesiredindexingdatatobeaccessedfromdiskduringretrieval.ThisworkissupportedbytheNationalScienceFoundationCenterforIntelligentInformationRetrievalattheUniversityofMassachusetts.Email:brown@cs.umass.edu.1IntroductionThehumanracehasalonghistoryofputtingthingsinwriting.Inoureffortstocommunicate,record,anddocument,weinevitablygeneratesomewrittenpieceofworkthatcontainshistory,facts,plans,orinformationingeneral.Whilegeneratingtheseworks,whichIshallcalldocuments,isrelativelyeasy,makinguseoftheinformationcontainedinthosedocumentscanbequitedifﬁcult,particularlyiftherearemanydocumentsinexistenceandweareinterestedinonlyafew.Itisexactlythisfunctionality,however,thatismostcritical.Theattorneymustbeabletoidentifypastcasesrelevanttothecurrentcase.Thescientistmustbeabletoidentifytechnicalpapersrelatedtotheirresearch.Theaccountantmustbeabletolocateapplicabletaxlaws.Thecontractormustbeabletoﬁndstandardsandspeciﬁcationsrelatedtotheproject.Anyoneinterestedinacurrenteventmustbeabletoﬁndrelevantnewsarticles.Awealthofdocumentsisuselessifwecannotidentifytheonesthatcontaintheinformationweneed.Oneoftheﬁrstsolutionstothisproblemappearednearlyfourthousandyearsagowhencataloguesofdocumentsinlibrarieswerecreatedtoaidinkeepingtrackofthosedocuments.Morerecently,inthe16thcentury,primitiveindexesfordocumentswerecreated.Anindexisalistofcertainkeywordsortopics.Eachentryinthelistcontainspointersintothedocumentswheredescriptionsanddiscussionsoftherespectivekeywordortopicmaybefound.Decidingwhatkeywordsandtopicsshouldgointoanindexandwhichdiscussionsareworthyofanindexpointerisatediousandsubjectivehumantaskpronetoomissions.Ashortcomingofanindexorcatalogueisthataninformationsearchmustbebasedoneitherthekeywordsactuallyindexedortheattributesusedinthecatalogue(e.g.,author,title,subject).Toaddressthislimitation,concordanceswerecreated.Aconcordanceforacollectionofdocumentscontainsanentryforeverytermthatappearsinthecollection.Aterm’sentrycontainsapointertoeveryoccurrenceofthatterm.Now,aninformationsearchwasnolongerrestrictedbypredeterminedkeywordsorattributes.However,aconcordanceforalargedocumentsuchastheBiblemightrequireagoodportionofalifetimetoconstructbyhand,andinonecase,theeffortinvolvedinsuchataskisbelievedtohaveleadtotheinsanityoftheconcordancecompiler[58].Withtheadventofthecomputerageinthelatterhalfofthe20thcentury,concordanceconstructioncouldbeautomated,greatlysimplifyingthetask.Whatusedtotakeyearscouldnowbeaccomplishedinminutes.Inspiteofbeingrelativelycompleteandsimpletoconstruct,aconcordancestillprovidedaratherunsophisticatedsolutiontoouroriginalproblem.Tryingtolocateinformationinalargecollectionofdocumentsusingaconcordancecanbeanexerciseinfrustration,leadingtotheretrievalofmanyunrelateddocumentsthatjusthappentocontaintermsthatwebelieveareindicativeoftheinformationweseek.Amoreintelligentsolutiontotheproblemathandwasstillneeded.Overthirtyyearsago,worktowardsthisintelligentsolutionbeganwiththebirthofinformationretrievalsystems.Thetaskofaninformationretrieval(IR)systemistosatisfyauser’sinformationneedbyidentifyingthedocumentsinacollectionofdocumentsthatcontainthedesiredinformation.Thistaskisaccomplishedthroughtheuseofthreebasiccomponents:adocumentrepresentation,aqueryrepresentation,andameasureofsimilaritybetweenqueriesanddocuments.Thedocumentrepresentationprovidesaformaldescriptionoftheinformationcontainedinthedocuments,thequeryrepresentationprovidesaformaldescriptionoftheinformationneed,andthesimilaritymeasuredeﬁnestherulesandproceduresformatchingtheinformationneedwiththedocumentsthatsatisfythatneed.MuchoftheresearchtodateinIRsystemshasfocussedonhowtoimplementthesethreecomponentsforbestperformance.IntheIRcommunity,“performance”isunderstoodtomeanretrievaleffectiveness,ratherthanexecutionefﬁciency.RetrievaleffectivenessreferstoanIRsystem’sabilitytocorrectlyidentifythedocumentsthatarerelevanttoagivenquery,andismeasuredintermsofrecallandprecision.Recallisthepercentageoftherelevantdocumentsactuallyidentiﬁedbythesystem,andprecisionisthepercentageofthedocumentsidentiﬁedthatarerelevant.F