Chapter 4 An Introduction to Hidden Markov Models

整理文档很辛苦,赏杯茶钱您下走!

免费阅读已结束,点击下载阅读编辑剩下 ...

阅读已结束,您可以下载文档离线阅读编辑

资源描述

Chapter4AnIntroductiontoHiddenMarkovModelsforBiologicalSequencesbyAndersKroghCenterforBiologicalSequenceAnalysisTechnicalUniversityofDenmarkBuilding206,2800Lyngby,DenmarkPhone:+4545252471Fax:+4545934808E-mail:krogh@cbs.dtu.dkInComputationalMethodsinMolecularBiology,editedbyS.L.Salzberg,D.B.SearlsandS.Kasif,pages45-63.Elsevier,1998.1Contents4AnIntroductiontoHiddenMarkovModelsforBiologicalSequences14.1Introduction.............................34.2FromregularexpressionstoHMMs................44.3ProfileHMMs............................84.3.1Pseudocounts........................104.3.2Searchingadatabase....................114.3.3Modelestimation......................124.4HMMsforgenefinding.......................134.4.1Signalsensors.......................144.4.2Codingregions.......................154.4.3Combiningthemodels...................164.5Furtherreading...........................1924.1IntroductionVeryefficientprogramsforsearchingatextforacombinationofwordsareavail-ableonmanycomputers.Thesamemethodscanbeusedforsearchingforpatternsinbiologicalsequences,butoftentheyfail.Thisisbecausebiological‘spelling’ismuchmoresloppythanEnglishspelling:proteinswiththesamefunctionfromtwodifferentorganismsarealmostcertainlyspelleddifferently,thatis,thetwoaminoacidsequencesdiffer.Itisnotrarethattwosuchhomologoussequenceshavelessthan30%identicalaminoacids.SimilarlyinDNAmanyinterestingsig-nalsvarygreatlyevenwithinthesamegenome.Somewell-knownexamplesareribosomebindingsitesandsplicesites,butthelistislong.Fortunatelythereareusuallystillsomesubtlesimilaritiesbetweentwosuchsequences,andtheques-tionishowtodetectthesesimilarities.Thevariationinafamilyofsequencescanbedescribedstatistically,andthisisthebasisformostmethodsusedinbiologicalsequenceanalysis,see[1]forapresentationofsomeofthesestatisticalapproaches.Forpairwisealignments,forinstance,theprobabilitythatacertainresiduemutatestoanotherresidueisusedinasubstitutionmatrix,suchasoneofthePAMmatrices.ForfindingpatternsinDNA,e.g.splicesites,somesortofweightmatrixisveryoftenused,whichissimplyapositionspecificscorecalculatedfromthefrequenciesofthefournucleotidesatallthepositionsinsomeknownexamples.Similarly,methodsforfindinggenesuse,almostwithoutexception,thestatisticsofcodonsordicodonsinsomeformorother.AhiddenMarkovmodel(HMM)isastatisticalmodel,whichisverywellsuitedformanytasksinmolecularbiology,althoughtheyhavebeenmostlyde-velopedforspeechrecognitionsincetheearly1970s,see[2]forhistoricaldetails.ThemostpopularuseoftheHMMinmolecularbiologyisasa‘probabilisticpro-file’ofaproteinfamily,whichiscalledaprofileHMM.Fromafamilyofproteins(orDNA)aprofileHMMcanbemadeforsearchingadatabaseforothermem-bersofthefamily.TheseprofileHMMsresembletheprofile[3]andweightmatrixmethods[4,5],andprobablythemaincontributionisthattheprofileHMMtreatsgapsinasystematicway.TheHMMcanbeappliedtoothertypesofproblems.Itisparticularlywellsuitedforproblemswithasimple‘grammaticalstructure,’suchasgenefinding.Ingenefindingseveralsignalsmustberecognizedandcombinedintoapredictionofexonsandintrons,andthepredictionmustconformtovariousrulestomakeitareasonablegeneprediction.AnHMMcancombinerecognitionofthesignals,anditcanbemadesuchthatthepredictionsalwaysfollowtherulesofagene.SincemuchoftheliteratureonHMMsisalittlehardtoreadformanybiol-ogists,Iwillattemptinthischaptertogiveanon-mathematicalintroductiontoHMMs.Whereasthelittlebiologicalbackgroundneededistakenforgranted,I3havetriedtoexplainHMMsatalevelthatalmostanyonecanfollow.FirstHMMsareintroducedbyanexampleandthenprofileHMMsaredescribed.ThenanHMMforfindingeukaryoticgenesissketched,andfinallypointerstothelitera-turearegiven.4.2FromregularexpressionstoHMMsMostreadershavenodoubtcomeacrossregularexpressionsatsomepoint,andmanyprobablyusethemquitealot.Regularexpressionsareusedinmanypro-grams,inparticularonUnixcomputers.Inprogramslikeawk,grep,sed,andperl,regularexpressionscanbeusedforsearchingtextfilesforapattern.Withgrepforinstance,youcansearchafileforalllinescontaining‘C.elegans’or‘Caenorhab-ditiselegans’withtheregularexpression‘’.ThiswillmatchanylinecontainingaCfollowedbyanynumberoflower-caselettersor‘.’,thenaspaceandthenelegans.Regularexpressionscanalsobeusedtocharacterizeproteinfamilies,whichisthebasisforthePROSITEdatabase[6].Usingregularexpressionsisaveryelegantandefficientwaytosearchforsomeproteinfamilies,butdifficultforother.Asalreadymentionedinthein-troduction,thedifficultiesarisebecauseproteinspellingismuchmorefreethanEnglishspelling.Thereforetheregularexpressionssometimesneedtobeverybroadandcomplex.ImagineaDNAmotiflikethis:!$#!$!!$##!$#!$(IuseDNAonlybecauseofthesmallernumberoflettersthanforaminoacids).Aregularexpressionforthisis[AT][CG][AC][ACGT]*A[TG][GC],meaningthatthefirstpositionisAorT,thesecondCorG,andsoforth.Theterm‘[ACGT]*’meansthatanyofthefourletterscanoccuranynumberoftimes.Theproblemwiththeaboveregularexpressionisthatitdoesnotinanywaydistinguishbetweenthehighlyimplausiblesequence!#%!##whichhastheexception

1 / 24
下载文档,编辑使用

©2015-2020 m.777doc.com 三七文档.

备案号:鲁ICP备2024069028号-1 客服联系 QQ:2149211541

×
保存成功