An Approach for Improving Execution Performance in

整理文档很辛苦,赏杯茶钱您下走!

免费阅读已结束,点击下载阅读编辑剩下 ...

阅读已结束,您可以下载文档离线阅读编辑

资源描述

AnApproachforImprovingExecutionPerformanceinInferenceNetworkBasedInformationRetrievalEricW.BrownTechnicalReport94–73September1994DepartmentofComputerScienceUniversityofMassachusettsAmherst,MA01003USAAbstractTheinferencenetworkretrievalmodelprovidestheabilitytocombineavarietyofretrievalstrategiesexpressedinarichquerylanguage.Whilethispoweryieldsimpressiveretrievaleffectiveness,italsopresentsbarrierstotheincorporationoftraditionaloptimizationtechniquesintendedtoimprovetheexecutionefficiency,orspeed,ofretrieval.Theessenceoftheseoptimizationtechniquesistheabilitytoidentifytheindexinginformationthatwillbemostusefulforevaluatingaquery,combinedwiththeabilitytoaccessfromdiskjustthatinformation.Isurveyavarietyoftechniquesintendedtoidentifyusefulindexinginformationandproposeanewstrategyforincorporatingthesetechniquesintotheinferencenetworkretrievalmodel.Additionally,Idescribeanarchitecturebasedonapersistentobjectstorethatprovidessophisticatedmanagementoftheindexingdata,allowingjustthedesiredindexingdatatobeaccessedfromdiskduringretrieval.ThisworkissupportedbytheNationalScienceFoundationCenterforIntelligentInformationRetrievalattheUniversityofMassachusetts.Email:brown@cs.umass.edu.1IntroductionThehumanracehasalonghistoryofputtingthingsinwriting.Inoureffortstocommunicate,record,anddocument,weinevitablygeneratesomewrittenpieceofworkthatcontainshistory,facts,plans,orinformationingeneral.Whilegeneratingtheseworks,whichIshallcalldocuments,isrelativelyeasy,makinguseoftheinformationcontainedinthosedocumentscanbequitedifficult,particularlyiftherearemanydocumentsinexistenceandweareinterestedinonlyafew.Itisexactlythisfunctionality,however,thatismostcritical.Theattorneymustbeabletoidentifypastcasesrelevanttothecurrentcase.Thescientistmustbeabletoidentifytechnicalpapersrelatedtotheirresearch.Theaccountantmustbeabletolocateapplicabletaxlaws.Thecontractormustbeabletofindstandardsandspecificationsrelatedtotheproject.Anyoneinterestedinacurrenteventmustbeabletofindrelevantnewsarticles.Awealthofdocumentsisuselessifwecannotidentifytheonesthatcontaintheinformationweneed.Oneofthefirstsolutionstothisproblemappearednearlyfourthousandyearsagowhencataloguesofdocumentsinlibrarieswerecreatedtoaidinkeepingtrackofthosedocuments.Morerecently,inthe16thcentury,primitiveindexesfordocumentswerecreated.Anindexisalistofcertainkeywordsortopics.Eachentryinthelistcontainspointersintothedocumentswheredescriptionsanddiscussionsoftherespectivekeywordortopicmaybefound.Decidingwhatkeywordsandtopicsshouldgointoanindexandwhichdiscussionsareworthyofanindexpointerisatediousandsubjectivehumantaskpronetoomissions.Ashortcomingofanindexorcatalogueisthataninformationsearchmustbebasedoneitherthekeywordsactuallyindexedortheattributesusedinthecatalogue(e.g.,author,title,subject).Toaddressthislimitation,concordanceswerecreated.Aconcordanceforacollectionofdocumentscontainsanentryforeverytermthatappearsinthecollection.Aterm’sentrycontainsapointertoeveryoccurrenceofthatterm.Now,aninformationsearchwasnolongerrestrictedbypredeterminedkeywordsorattributes.However,aconcordanceforalargedocumentsuchastheBiblemightrequireagoodportionofalifetimetoconstructbyhand,andinonecase,theeffortinvolvedinsuchataskisbelievedtohaveleadtotheinsanityoftheconcordancecompiler[58].Withtheadventofthecomputerageinthelatterhalfofthe20thcentury,concordanceconstructioncouldbeautomated,greatlysimplifyingthetask.Whatusedtotakeyearscouldnowbeaccomplishedinminutes.Inspiteofbeingrelativelycompleteandsimpletoconstruct,aconcordancestillprovidedaratherunsophisticatedsolutiontoouroriginalproblem.Tryingtolocateinformationinalargecollectionofdocumentsusingaconcordancecanbeanexerciseinfrustration,leadingtotheretrievalofmanyunrelateddocumentsthatjusthappentocontaintermsthatwebelieveareindicativeoftheinformationweseek.Amoreintelligentsolutiontotheproblemathandwasstillneeded.Overthirtyyearsago,worktowardsthisintelligentsolutionbeganwiththebirthofinformationretrievalsystems.Thetaskofaninformationretrieval(IR)systemistosatisfyauser’sinformationneedbyidentifyingthedocumentsinacollectionofdocumentsthatcontainthedesiredinformation.Thistaskisaccomplishedthroughtheuseofthreebasiccomponents:adocumentrepresentation,aqueryrepresentation,andameasureofsimilaritybetweenqueriesanddocuments.Thedocumentrepresentationprovidesaformaldescriptionoftheinformationcontainedinthedocuments,thequeryrepresentationprovidesaformaldescriptionoftheinformationneed,andthesimilaritymeasuredefinestherulesandproceduresformatchingtheinformationneedwiththedocumentsthatsatisfythatneed.MuchoftheresearchtodateinIRsystemshasfocussedonhowtoimplementthesethreecomponentsforbestperformance.IntheIRcommunity,“performance”isunderstoodtomeanretrievaleffectiveness,ratherthanexecutionefficiency.RetrievaleffectivenessreferstoanIRsystem’sabilitytocorrectlyidentifythedocumentsthatarerelevanttoagivenquery,andismeasuredintermsofrecallandprecision.Recallisthepercentageoftherelevantdocumentsactuallyidentifiedbythesystem,andprecisionisthepercentageofthedocumentsidentifiedthatarerelevant.F

1 / 33
下载文档,编辑使用

©2015-2020 m.777doc.com 三七文档.

备案号:鲁ICP备2024069028号-1 客服联系 QQ:2149211541

×
保存成功