NCBI Home Page NCBI Site Search page NCBI Guide that lists and describes the NCBI resources
Conserved domains on  [gi|568954801|ref|XP_006509472|]
View 

teneurin-3 isoform X15 [Mus musculus]

Protein Classification

Graphical summary

 Zoom to residue level

show extra options »

Show site features     Horizontal zoom: ×

List of domain hits

Name Accession Description Interval E-value
NHL super family cl18310
NHL repeat unit of beta-propeller proteins; The NHL(NCL-1, HT2A and LIN-41)-repeat is found in ...
890-1249 2.53e-41

NHL repeat unit of beta-propeller proteins; The NHL(NCL-1, HT2A and LIN-41)-repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures. The repeats have a catalytic activity in Peptidyl-glycine alpha-amidating monooxygenase; proteolysis has shown that the Peptidyl-alpha-hydroxyglycine alpha-amidating lyase (PAL) activity is localized to the repeats. Tripartite motif-containing protein 32 interacts with the activation domain of Tat. This interaction is mediated by the NHL repeats.


The actual alignment was detected with superfamily member cd14953:

Pssm-ID: 302697 [Multi-domain]  Cd Length: 323  Bit Score: 156.15  E-value: 2.53e-41
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  890 VSSIMGNGRRrsiscpscnGQADGNKLLA----PVALACGIDGSLYVGDF--NYVRRIFPSGNVTSVLELRNKDFRHSSN 963
Cdd:cd14953     1 VSTVAGSGTA---------GFSGGGGTAArfnsPSGVAVDAAGNLYVADRgnHRIRKITPDGVVTTVAGTGTAGFADGGG 71
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  964 PAHRYY----LATDPvTGDLYVSDTNTRRIYRpksLTGAKDLTknaeVVAGTGEqclpfdeARCGDGGKAVEATLMSPKG 1039
Cdd:cd14953    72 AAAQFNtpsgVAVDA-AGNLYVADTGNHRIRK---ITPDGVVS----TLAGTGT-------AGFSDDGGATAAQFNYPTG 136
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1040 MAIDKNGLIYFVDGT--MIRKVDQNGIISTLLGSNDLTSAR--PLTcdtsmhisQVRLEWPTDLAINPMDNsIYVLD--N 1113
Cdd:cd14953   137 VAVDAAGNLYVADTGnhRIRKITPDGVVTTVAGTGGAGYAGdgPAT--------AAQFNNPTGVAVDAAGN-LYVADrgN 207
                         250       260       270       280       290       300       310       320
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1114 NVVLQITENRQVRIAAGRPmhcqvpGVEYPVGKHAVQTTLESATAIAVSYSGVLYITETDEkkiNRIRQVTTDGEISLVA 1193
Cdd:cd14953   208 HRIRKITPDGVVTTVAGTG------TAGFSGDGGATAAQLNNPTGVAVDAAGNLYVADSGN---HRIRKITPAGVVTTVA 278
                         330       340       350       360       370
                  ....*....|....*....|....*....|....*....|....*....|....*..
gi 568954801 1194 GIPSEcdckndancdcyQSGD-GYAKDAKLNAPSSLAASPDGTLYIADLGNIRIRAV 1249
Cdd:cd14953   279 GGGAG------------FSGDgGPATSAQFNNPTGVAVDAAGNLYVADTGNNRIRKI 323
Tox-GHH pfam15636
GHH signature containing HNH/Endo VII superfamily nuclease toxin; A predicted toxin of the HNH ...
2366-2443 7.00e-40

GHH signature containing HNH/Endo VII superfamily nuclease toxin; A predicted toxin of the HNH/Endonuclease VII fold present in bacterial polymorphic toxin systems with a characteriztic sG[HQ]H signature motif. In bacterial polymorphic toxin systems, the toxin is exported by the type 2, type 6, type 7 or TcdB/TcaC-type secretion system. The metazoan teneurin proteins possess an inactive of this domain at their C-terminus.


:

Pssm-ID: 464783  Cd Length: 78  Bit Score: 142.75  E-value: 7.00e-40
                           10        20        30        40        50        60        70
                   ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*...
gi 568954801  2366 EEKARILEQARQRALARAWAREQQRVRDGEEGARLWTEGEKRQLLSAGKVQGYDGYYVLSVEQYPELADSANNIQFLR 2443
Cdd:pfam15636    1 EERKRLLEHAKKRAVREAWHRERQLLRNGLPGSRDWTDEEKEELLSTGSVPGYDGEYIHPVEQYPELADDPSNIRFRK 78
RhsA COG3209
Uncharacterized conserved protein RhaS, contains 28 RHS repeats [General function prediction ...
1203-2143 4.86e-34

Uncharacterized conserved protein RhaS, contains 28 RHS repeats [General function prediction only];


:

Pssm-ID: 442442 [Multi-domain]  Cd Length: 1103  Bit Score: 144.13  E-value: 4.86e-34
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1203 NDANCDCYQSGDGYAKDAKLNAPSSLAASPDGTLYIADLGNIRIRAVSKNKPLLNSMNFYEVASPTDQELYIFDINGTHQ 1282
Cdd:COG3209   109 AAATASAGRLVSTGAGAGGTVTAATGGTLGATAGSATTGSTDGGRGGVAVTGLAGGGASAYGLTLGGAAAGPATGVGTGA 188
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1283 YTVSLVTGDYLYNFSYSNDNDVTAVTDSNGNTLRIRRDPNRMPVRVVSPDNQVIWLTIGTNGCLKSMTAQGLELVLFTYH 1362
Cdd:COG3209   189 VTLATGLAGSALLALGSGAILGGLAGAYSGSATTATGTALGTPASVAATVTGSATGAAGAGAAVATAATTLGGTTGAGTG 268
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1363 GNSGLLATK---SDETGWTTFFDYDSEGRLTNVTFPTGVVTNLHGDMDKAITVDIESSSREEDVSITSNLSSIDSFYTMV 1439
Cdd:COG3209   269 ASGAGLDAStgtGGAGGSNAAATAGGLGGAGLGSGGAGGGGTAGGTTTAAGTTGTAAVSGAADAGTTTTTGTGTGGTTTT 348
                         250       260       270       280       290       300       310       320
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1440 QDQLRNSYQIGYDGSLRIFYASGLDSHYQTEPHVLAGTANPTVAKRNMTLPGENGQNLVEWRFRKEQAQGKVNVFGRKLR 1519
Cdd:COG3209   349 VGGGGSLTLGGYGAAGGLTTSVGAGGGGSTSGSTTTVGGGGTATGSGGGSSTTGVGAGTTTTSTTGGDGGPATAAGALTA 428
                         330       340       350       360       370       380       390       400
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1520 VNGRNLLSVDFDRTTKTEKIYDDHRKFLLRIAYDTSGHPTLWLPSSKLMAVNVTYSSTGQIASIQRGTTSEKVDYDSQGR 1599
Cdd:COG3209   429 GGTATGTGTGGGGTTAGTDATTTTGGAGASGTLTTTGGAATGATTGGGTEAGTGGGTLTSGSAGATTLGTDTTLDDTLGG 508
                         410       420       430       440       450       460       470       480
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1600 IVSRVFADGKTWSYTYLEKSMVLLLHSQRQYIFEYDMWDRLSAITMPSVARHTMQTIRSIGYYRNIYNPPESNASIITDY 1679
Cdd:COG3209   509 TTTTTAGARGLVVTTGTTLTLGTTTTATLSATDATGTGDTTTTGTVGTGTSTGTGGTGTVTTTGDGTGGASTTTGTTGGT 588
                         490       500       510       520       530       540       550       560
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1680 NEEGLLLQTAFLGTSRRVLFKYRRQTRLSEILYDSTRVSFTYDETAGVLKTVNLQSDGFicTIRYRQIGPLIDRQIFRFS 1759
Cdd:COG3209   589 ATTTTVTTTTTTSTAGTTTTTTSGYTRAGLTLTLGTGTASGLERATASTGSTTGGTTGT--GVTTTGTTTTRATGTTGTG 666
                         570       580       590       600       610       620       630       640
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1760 EDGMVNARFDYSYDNSFRVTSMQGVINETPLPIDLYQFDDISGKVEQFGKFGVIYYDINQIISTAVMTYTKHFDAHGRIK 1839
Cdd:COG3209   667 TGVTAGLTTLATGGTTVGGGTGTTSTATTGATTGGTETGTTVTTLAGGTTTRLGTTTTGGGGGTTTDGTGTGGTTGTLTT 746
                         650       660       670       680       690       700       710       720
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1840 EIQYEifRSLMYWITIQYDNMGRVTKREIKIGPFANTTKYAYEYDVDGQLQTVYLNEKIMWRYNYDLNGNLHLLNPSSSA 1919
Cdd:COG3209   747 TSTTT--TTTAGALTYTYDALGRLTSETTPGGVTQGTYTTRYTYDALGRLTSVTYPDGETVTYTYDALGRLTSVITVGSG 824
                         730       740       750       760       770       780       790       800
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1920 RLTPL-----RYDLRDRITRLgdvqyrldEDGFLRQRGTEIFEYSSKGLLTRVYSKGSGWTviYRYDGLGRRVSSKTSLG 1994
Cdd:COG3209   825 GGTDLqdrtyTYDAAGNITSI--------TDALRAGTLTQTYTYDALGRLTSATDPGTTES--YTYDANGNLTSRTDGGT 894
                         810       820       830       840       850       860       870       880
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1995 QHLQFFYADLtyPTRITHvynhSSSEITSLYYDLQGHlfameissgdefyiaSDNTGTPLAVFSSNGLMLKQIQYTAYGE 2074
Cdd:COG3209   895 TTYTYDALGR--LVSVTK----PDGTTTTYTYDALGH---------------TDHLGSVRALTDASGQVVWRYDYDPFGN 953
                         890       900       910       920       930       940
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....
gi 568954801 2075 IYFDSNVDFQLVIGFHGGLYDPLTKLIHFGERDYDILAGRWTTPDieiwkRIGKDPAPfNLYMFRNNNP 2143
Cdd:COG3209   954 LLAETSGAAANPLRFTGQEYDAETGLYYNGARYYDPALGRFLSPD-----PIGLAGGL-NLYAYVGNNP 1016
Ten_N super family cl24184
Teneurin Intracellular Region; This family is found in the intracellular N-terminal region of ...
1-36 1.13e-17

Teneurin Intracellular Region; This family is found in the intracellular N-terminal region of the Teneurin family of proteins. These proteins are 'pair-rule' genes and are involved in tissue patterning, specifically probably neural patterning. The intracellular domain is cleaved in response to homophilic interaction of the extracellular domain, and translocates to the nucleus. Here it probably carries out to some transcriptional regulatory activity. The length of this region and the conservation suggests that there may be two structural domains here (personal obs:C Yeats).


The actual alignment was detected with superfamily member pfam06484:

Pssm-ID: 461932 [Multi-domain]  Cd Length: 367  Bit Score: 87.34  E-value: 1.13e-17
                           10        20        30
                   ....*....|....*....|....*....|....*.
gi 568954801     1 MASGSVYSPPTRPLPRNTLSRSAFKFKKSSKYCSWR 36
Cdd:pfam06484  332 LTSGTVYSPPPRPLPRNTFSRPAFKLKKPYKYCSWK 367
DUF5885 super family cl44670
Family of unknown function (DUF5885); This is a family of uncharacterized proteins of unknown ...
250-435 1.99e-08

Family of unknown function (DUF5885); This is a family of uncharacterized proteins of unknown function found in viruses.


The actual alignment was detected with superfamily member pfam19232:

Pssm-ID: 437064  Cd Length: 265  Bit Score: 57.71  E-value: 1.99e-08
                           10        20        30        40        50        60        70        80
                   ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801   250 CHGNGECvsGTCH-CFPGFLGPDCSraACPVLCSGNGQ----------YSKGRC----LCFSGwkgTECDVPTTQCI-DP 313
Cdd:pfam19232   34 CTTDAQC--GTCMtCVAGACTPKAS--CCGGVTCGAGQtcdaktntcvYVKGYCsadhPCPSG---SACDTAKNACIaQP 106
                           90       100       110       120       130       140       150       160
                   ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801   314 QCG---GRGiCIMG-------------------SCACNSGYK-GENCE--------EADCLDP---------------GC 347
Cdd:pfam19232  107 PYGpdsGKG-CVRGfgawiweldpatnsgvwrcRCANGSLYNsAHECSpladqtlcAAENLDPnalvpassvpafaayGW 185
                          170       180       190       200       210       220       230       240
                   ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801   348 SNHGVCIH-------------GECHCNPGWGGSNCEILKTmcadqCSGHGTYLQESGSCTCDPNWTGpdcsneicsvdcg 414
Cdd:pfam19232  186 GNQPVLINkstagaavpsplaGVCPCKPGWAGGSCTEDRT-----CNGRGTWNETTGQCACNIDFSG------------- 247
                          250       260
                   ....*....|....*....|....
gi 568954801   415 shgvcmGGSCRCEEG---WTGPAC 435
Cdd:pfam19232  248 ------HNSCGDDNNctsWTGPRC 265
acid_disulf_rpt NF033662
acidic double-disulfide repeat; The acidic double-disulfide repeat is an Asp-rich repeat with ...
520-550 9.90e-08

acidic double-disulfide repeat; The acidic double-disulfide repeat is an Asp-rich repeat with four nearly invariant Cys residues in a repeat length of about 35 amino acids.


:

Pssm-ID: 411265 [Multi-domain]  Cd Length: 32  Bit Score: 49.82  E-value: 9.90e-08
                          10        20        30
                  ....*....|....*....|....*....|.
gi 568954801  520 AMETLCTDSKDNEGDGLIDCMDPDCCLQSSC 550
Cdd:NF033662    2 ATDTTCSDGIDNDGDGLTDCADPDCAGNPVC 32
DSL super family cl19567
Delta serrate ligand; This family has been redefined to correspond to the EGF-like domain ...
426-469 1.61e-05

Delta serrate ligand; This family has been redefined to correspond to the EGF-like domain defined by structure.


The actual alignment was detected with superfamily member pfam01414:

Pssm-ID: 473190  Cd Length: 46  Bit Score: 44.15  E-value: 1.61e-05
                           10        20        30        40
                   ....*....|....*....|....*....|....*....|....*..
gi 568954801   426 CEEGWTGPACNqRACHPRCAE--HGTC-KDGKCECSQGWNGEHCTIA 469
Cdd:pfam01414    1 CDENYYGSTCS-KFCRPRDDKfgHYTCdANGNKVCLPGWTGPYCDKP 46
EGF_CA cd00054
Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular ...
488-517 3.62e-03

Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular (mostly animal) proteins. Many of these proteins require calcium for their biological function and calcium-binding sites have been found to be located at the N-terminus of particular EGF-like domains; calcium-binding may be crucial for numerous protein-protein interactions. Six conserved core cysteines form three disulfide bridges as in non calcium-binding EGF domains, whose structures are very similar. EGF_CA can be found in tandem repeat arrangements.


:

Pssm-ID: 238011  Cd Length: 38  Bit Score: 37.23  E-value: 3.62e-03
                          10        20        30
                  ....*....|....*....|....*....|
gi 568954801  488 PGLCNSNGRCTLDQNGWHCVCQPGWRGAGC 517
Cdd:cd00054     8 GNPCQNGGTCVNTVGSYRCSCPPGYTGRNC 37
 
Name Accession Description Interval E-value
NHL_like_1 cd14953
Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat ...
890-1249 2.53e-41

Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat domains is found in a variety of domain architectures. The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.


Pssm-ID: 271323 [Multi-domain]  Cd Length: 323  Bit Score: 156.15  E-value: 2.53e-41
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  890 VSSIMGNGRRrsiscpscnGQADGNKLLA----PVALACGIDGSLYVGDF--NYVRRIFPSGNVTSVLELRNKDFRHSSN 963
Cdd:cd14953     1 VSTVAGSGTA---------GFSGGGGTAArfnsPSGVAVDAAGNLYVADRgnHRIRKITPDGVVTTVAGTGTAGFADGGG 71
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  964 PAHRYY----LATDPvTGDLYVSDTNTRRIYRpksLTGAKDLTknaeVVAGTGEqclpfdeARCGDGGKAVEATLMSPKG 1039
Cdd:cd14953    72 AAAQFNtpsgVAVDA-AGNLYVADTGNHRIRK---ITPDGVVS----TLAGTGT-------AGFSDDGGATAAQFNYPTG 136
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1040 MAIDKNGLIYFVDGT--MIRKVDQNGIISTLLGSNDLTSAR--PLTcdtsmhisQVRLEWPTDLAINPMDNsIYVLD--N 1113
Cdd:cd14953   137 VAVDAAGNLYVADTGnhRIRKITPDGVVTTVAGTGGAGYAGdgPAT--------AAQFNNPTGVAVDAAGN-LYVADrgN 207
                         250       260       270       280       290       300       310       320
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1114 NVVLQITENRQVRIAAGRPmhcqvpGVEYPVGKHAVQTTLESATAIAVSYSGVLYITETDEkkiNRIRQVTTDGEISLVA 1193
Cdd:cd14953   208 HRIRKITPDGVVTTVAGTG------TAGFSGDGGATAAQLNNPTGVAVDAAGNLYVADSGN---HRIRKITPAGVVTTVA 278
                         330       340       350       360       370
                  ....*....|....*....|....*....|....*....|....*....|....*..
gi 568954801 1194 GIPSEcdckndancdcyQSGD-GYAKDAKLNAPSSLAASPDGTLYIADLGNIRIRAV 1249
Cdd:cd14953   279 GGGAG------------FSGDgGPATSAQFNNPTGVAVDAAGNLYVADTGNNRIRKI 323
Tox-GHH pfam15636
GHH signature containing HNH/Endo VII superfamily nuclease toxin; A predicted toxin of the HNH ...
2366-2443 7.00e-40

GHH signature containing HNH/Endo VII superfamily nuclease toxin; A predicted toxin of the HNH/Endonuclease VII fold present in bacterial polymorphic toxin systems with a characteriztic sG[HQ]H signature motif. In bacterial polymorphic toxin systems, the toxin is exported by the type 2, type 6, type 7 or TcdB/TcaC-type secretion system. The metazoan teneurin proteins possess an inactive of this domain at their C-terminus.


Pssm-ID: 464783  Cd Length: 78  Bit Score: 142.75  E-value: 7.00e-40
                           10        20        30        40        50        60        70
                   ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*...
gi 568954801  2366 EEKARILEQARQRALARAWAREQQRVRDGEEGARLWTEGEKRQLLSAGKVQGYDGYYVLSVEQYPELADSANNIQFLR 2443
Cdd:pfam15636    1 EERKRLLEHAKKRAVREAWHRERQLLRNGLPGSRDWTDEEKEELLSTGSVPGYDGEYIHPVEQYPELADDPSNIRFRK 78
RhsA COG3209
Uncharacterized conserved protein RhaS, contains 28 RHS repeats [General function prediction ...
1203-2143 4.86e-34

Uncharacterized conserved protein RhaS, contains 28 RHS repeats [General function prediction only];


Pssm-ID: 442442 [Multi-domain]  Cd Length: 1103  Bit Score: 144.13  E-value: 4.86e-34
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1203 NDANCDCYQSGDGYAKDAKLNAPSSLAASPDGTLYIADLGNIRIRAVSKNKPLLNSMNFYEVASPTDQELYIFDINGTHQ 1282
Cdd:COG3209   109 AAATASAGRLVSTGAGAGGTVTAATGGTLGATAGSATTGSTDGGRGGVAVTGLAGGGASAYGLTLGGAAAGPATGVGTGA 188
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1283 YTVSLVTGDYLYNFSYSNDNDVTAVTDSNGNTLRIRRDPNRMPVRVVSPDNQVIWLTIGTNGCLKSMTAQGLELVLFTYH 1362
Cdd:COG3209   189 VTLATGLAGSALLALGSGAILGGLAGAYSGSATTATGTALGTPASVAATVTGSATGAAGAGAAVATAATTLGGTTGAGTG 268
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1363 GNSGLLATK---SDETGWTTFFDYDSEGRLTNVTFPTGVVTNLHGDMDKAITVDIESSSREEDVSITSNLSSIDSFYTMV 1439
Cdd:COG3209   269 ASGAGLDAStgtGGAGGSNAAATAGGLGGAGLGSGGAGGGGTAGGTTTAAGTTGTAAVSGAADAGTTTTTGTGTGGTTTT 348
                         250       260       270       280       290       300       310       320
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1440 QDQLRNSYQIGYDGSLRIFYASGLDSHYQTEPHVLAGTANPTVAKRNMTLPGENGQNLVEWRFRKEQAQGKVNVFGRKLR 1519
Cdd:COG3209   349 VGGGGSLTLGGYGAAGGLTTSVGAGGGGSTSGSTTTVGGGGTATGSGGGSSTTGVGAGTTTTSTTGGDGGPATAAGALTA 428
                         330       340       350       360       370       380       390       400
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1520 VNGRNLLSVDFDRTTKTEKIYDDHRKFLLRIAYDTSGHPTLWLPSSKLMAVNVTYSSTGQIASIQRGTTSEKVDYDSQGR 1599
Cdd:COG3209   429 GGTATGTGTGGGGTTAGTDATTTTGGAGASGTLTTTGGAATGATTGGGTEAGTGGGTLTSGSAGATTLGTDTTLDDTLGG 508
                         410       420       430       440       450       460       470       480
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1600 IVSRVFADGKTWSYTYLEKSMVLLLHSQRQYIFEYDMWDRLSAITMPSVARHTMQTIRSIGYYRNIYNPPESNASIITDY 1679
Cdd:COG3209   509 TTTTTAGARGLVVTTGTTLTLGTTTTATLSATDATGTGDTTTTGTVGTGTSTGTGGTGTVTTTGDGTGGASTTTGTTGGT 588
                         490       500       510       520       530       540       550       560
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1680 NEEGLLLQTAFLGTSRRVLFKYRRQTRLSEILYDSTRVSFTYDETAGVLKTVNLQSDGFicTIRYRQIGPLIDRQIFRFS 1759
Cdd:COG3209   589 ATTTTVTTTTTTSTAGTTTTTTSGYTRAGLTLTLGTGTASGLERATASTGSTTGGTTGT--GVTTTGTTTTRATGTTGTG 666
                         570       580       590       600       610       620       630       640
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1760 EDGMVNARFDYSYDNSFRVTSMQGVINETPLPIDLYQFDDISGKVEQFGKFGVIYYDINQIISTAVMTYTKHFDAHGRIK 1839
Cdd:COG3209   667 TGVTAGLTTLATGGTTVGGGTGTTSTATTGATTGGTETGTTVTTLAGGTTTRLGTTTTGGGGGTTTDGTGTGGTTGTLTT 746
                         650       660       670       680       690       700       710       720
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1840 EIQYEifRSLMYWITIQYDNMGRVTKREIKIGPFANTTKYAYEYDVDGQLQTVYLNEKIMWRYNYDLNGNLHLLNPSSSA 1919
Cdd:COG3209   747 TSTTT--TTTAGALTYTYDALGRLTSETTPGGVTQGTYTTRYTYDALGRLTSVTYPDGETVTYTYDALGRLTSVITVGSG 824
                         730       740       750       760       770       780       790       800
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1920 RLTPL-----RYDLRDRITRLgdvqyrldEDGFLRQRGTEIFEYSSKGLLTRVYSKGSGWTviYRYDGLGRRVSSKTSLG 1994
Cdd:COG3209   825 GGTDLqdrtyTYDAAGNITSI--------TDALRAGTLTQTYTYDALGRLTSATDPGTTES--YTYDANGNLTSRTDGGT 894
                         810       820       830       840       850       860       870       880
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1995 QHLQFFYADLtyPTRITHvynhSSSEITSLYYDLQGHlfameissgdefyiaSDNTGTPLAVFSSNGLMLKQIQYTAYGE 2074
Cdd:COG3209   895 TTYTYDALGR--LVSVTK----PDGTTTTYTYDALGH---------------TDHLGSVRALTDASGQVVWRYDYDPFGN 953
                         890       900       910       920       930       940
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....
gi 568954801 2075 IYFDSNVDFQLVIGFHGGLYDPLTKLIHFGERDYDILAGRWTTPDieiwkRIGKDPAPfNLYMFRNNNP 2143
Cdd:COG3209   954 LLAETSGAAANPLRFTGQEYDAETGLYYNGARYYDPALGRFLSPD-----PIGLAGGL-NLYAYVGNNP 1016
Ten_N pfam06484
Teneurin Intracellular Region; This family is found in the intracellular N-terminal region of ...
1-36 1.13e-17

Teneurin Intracellular Region; This family is found in the intracellular N-terminal region of the Teneurin family of proteins. These proteins are 'pair-rule' genes and are involved in tissue patterning, specifically probably neural patterning. The intracellular domain is cleaved in response to homophilic interaction of the extracellular domain, and translocates to the nucleus. Here it probably carries out to some transcriptional regulatory activity. The length of this region and the conservation suggests that there may be two structural domains here (personal obs:C Yeats).


Pssm-ID: 461932 [Multi-domain]  Cd Length: 367  Bit Score: 87.34  E-value: 1.13e-17
                           10        20        30
                   ....*....|....*....|....*....|....*.
gi 568954801     1 MASGSVYSPPTRPLPRNTLSRSAFKFKKSSKYCSWR 36
Cdd:pfam06484  332 LTSGTVYSPPPRPLPRNTFSRPAFKLKKPYKYCSWK 367
Rhs_assc_core TIGR03696
RHS repeat-associated core domain; This model represents a conserved unique core sequence ...
2069-2143 3.40e-09

RHS repeat-associated core domain; This model represents a conserved unique core sequence shared by large numbers of proteins. It is occasional in the Archaea Methanosarcina barkeri) but common in bacteria and eukaryotes. Most fall into two large classes. One class consists of long proteins in which two classes of repeats are abundant: an FG-GAP repeat (pfam01839) class, and an RHS repeat (pfam05593) or YD repeat (TIGR01643). This class includes secreted bacterial insecticidal toxins and intercellular signalling proteins such as the teneurins in animals. The other class consists of uncharacterized proteins shorter than 400 amino acids, where this core domain of about 75 amino acids tends to occur in the N-terminal half. Over twenty such proteins are found in Pseudomonas putida alone; little sequence similarity or repeat structure is found among these proteins outside the region modeled by this domain.


Pssm-ID: 274730 [Multi-domain]  Cd Length: 77  Bit Score: 55.20  E-value: 3.40e-09
                           10        20        30        40        50        60        70
                   ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*
gi 568954801  2069 YTAYGEIYFDSNVDFQLvIGFHGGLYDPLTKLIHFGERDYDILAGRWTTPDieiwkRIGKDpAPFNLYMFRNNNP 2143
Cdd:TIGR03696    1 YDPYGEVLSESGAAPNP-LRFTGQYYDAETGLYYNGARYYDPELGRFLSPD-----PIGLG-GGLNLYAYVGNNP 68
Vgb COG4257
Streptogramin lyase [Defense mechanisms];
919-1197 3.43e-09

Streptogramin lyase [Defense mechanisms];


Pssm-ID: 443399 [Multi-domain]  Cd Length: 270  Bit Score: 60.42  E-value: 3.43e-09
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  919 PVALACGIDGSLYVGDF--NYVRRIFP-SGNVTsvlelrnkdfRHSSNPAHRYY-LATDPvTGDLYVSDTNTRRIYRpks 994
Cdd:COG4257    19 PRDVAVDPDGAVWFTDQggGRIGRLDPaTGEFT----------EYPLGGGSGPHgIAVDP-DGNLWFTDNGNNRIGR--- 84
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  995 LTGAkdlTKNAEVVAGTGEQCLPFdearcgdggkaveatlmspkGMAIDKNGLIYFVDGT--MIRKVD-QNGIISTLLGS 1071
Cdd:COG4257    85 IDPK---TGEITTFALPGGGSNPH--------------------GIAFDPDGNLWFTDQGgnRIGRLDpATGEVTEFPLP 141
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1072 NDLTSARPLTCD--------------------TSMHISQVRLE----WPTDLAINPmDNSIYVLD--NNVVLQITEnrqv 1125
Cdd:COG4257   142 TGGAGPYGIAVDpdgnlwvtdfganaigridpDTGTLTEYALPtpgaGPRGLAVDP-DGNLWVADtgSGRIGRFDP---- 216
                         250       260       270       280       290       300       310
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|..
gi 568954801 1126 riAAGRpmhcqvpgveypVGKHAVQTTLESATAIAVSYSGVLYITETDekkINRIRQVTTDGEISLVAgIPS 1197
Cdd:COG4257   217 --KTGT------------VTEYPLPGGGARPYGVAVDGDGRVWFAESG---ANRIVRFDPDTELTEYV-LPS 270
DUF5885 pfam19232
Family of unknown function (DUF5885); This is a family of uncharacterized proteins of unknown ...
250-435 1.99e-08

Family of unknown function (DUF5885); This is a family of uncharacterized proteins of unknown function found in viruses.


Pssm-ID: 437064  Cd Length: 265  Bit Score: 57.71  E-value: 1.99e-08
                           10        20        30        40        50        60        70        80
                   ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801   250 CHGNGECvsGTCH-CFPGFLGPDCSraACPVLCSGNGQ----------YSKGRC----LCFSGwkgTECDVPTTQCI-DP 313
Cdd:pfam19232   34 CTTDAQC--GTCMtCVAGACTPKAS--CCGGVTCGAGQtcdaktntcvYVKGYCsadhPCPSG---SACDTAKNACIaQP 106
                           90       100       110       120       130       140       150       160
                   ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801   314 QCG---GRGiCIMG-------------------SCACNSGYK-GENCE--------EADCLDP---------------GC 347
Cdd:pfam19232  107 PYGpdsGKG-CVRGfgawiweldpatnsgvwrcRCANGSLYNsAHECSpladqtlcAAENLDPnalvpassvpafaayGW 185
                          170       180       190       200       210       220       230       240
                   ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801   348 SNHGVCIH-------------GECHCNPGWGGSNCEILKTmcadqCSGHGTYLQESGSCTCDPNWTGpdcsneicsvdcg 414
Cdd:pfam19232  186 GNQPVLINkstagaavpsplaGVCPCKPGWAGGSCTEDRT-----CNGRGTWNETTGQCACNIDFSG------------- 247
                          250       260
                   ....*....|....*....|....
gi 568954801   415 shgvcmGGSCRCEEG---WTGPAC 435
Cdd:pfam19232  248 ------HNSCGDDNNctsWTGPRC 265
acid_disulf_rpt NF033662
acidic double-disulfide repeat; The acidic double-disulfide repeat is an Asp-rich repeat with ...
520-550 9.90e-08

acidic double-disulfide repeat; The acidic double-disulfide repeat is an Asp-rich repeat with four nearly invariant Cys residues in a repeat length of about 35 amino acids.


Pssm-ID: 411265 [Multi-domain]  Cd Length: 32  Bit Score: 49.82  E-value: 9.90e-08
                          10        20        30
                  ....*....|....*....|....*....|.
gi 568954801  520 AMETLCTDSKDNEGDGLIDCMDPDCCLQSSC 550
Cdd:NF033662    2 ATDTTCSDGIDNDGDGLTDCADPDCAGNPVC 32
C_rich_MXAN6577 NF041328
MXAN_6577-like cysteine-rich domain;
301-455 6.92e-07

MXAN_6577-like cysteine-rich domain;


Pssm-ID: 469225 [Multi-domain]  Cd Length: 145  Bit Score: 50.91  E-value: 6.92e-07
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  301 TECDVPTTQCIDPQ--CGGRGICIMGScACNSGykgeNCEEAdcldpgCSNHGVCIHGECHCNPGwggsnceilKTMCAD 378
Cdd:NF041328   12 AGCPEPGAVCPEGLsvCGGACVDLRSD-PSNCG----ACGVA------CGAGQTCVAGACGCGPG---------TVACGG 71
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  379 QCSGHGTylqesgsctcDPNWTGPdcsneiCSVDCGSHGVCMGGSCR--CEEGWT--GPAC--------NQRACHPRCAE 446
Cdd:NF041328   72 ACVDTAS----------DPAHCGA------CGAACAPGQVCEGGACReaCSEGLTrcGGACvdlatdplHCGACGVACDP 135

                  ....*....
gi 568954801  447 HGTCKDGKC 455
Cdd:NF041328  136 GESCRGGAC 144
DSL pfam01414
Delta serrate ligand; This family has been redefined to correspond to the EGF-like domain ...
426-469 1.61e-05

Delta serrate ligand; This family has been redefined to correspond to the EGF-like domain defined by structure.


Pssm-ID: 460202  Cd Length: 46  Bit Score: 44.15  E-value: 1.61e-05
                           10        20        30        40
                   ....*....|....*....|....*....|....*....|....*..
gi 568954801   426 CEEGWTGPACNqRACHPRCAE--HGTC-KDGKCECSQGWNGEHCTIA 469
Cdd:pfam01414    1 CDENYYGSTCS-KFCRPRDDKfgHYTCdANGNKVCLPGWTGPYCDKP 46
PLN02919 PLN02919
haloacid dehalogenase-like hydrolase family protein
970-1253 3.45e-05

haloacid dehalogenase-like hydrolase family protein


Pssm-ID: 215497 [Multi-domain]  Cd Length: 1057  Bit Score: 49.46  E-value: 3.45e-05
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  970 LATDPVTGDLYVSDTNTRRIYrpksltgAKDLTKNAEV-VAGTGEQCL---PFDEArcgdggkaveaTLMSPKGMAID-K 1044
Cdd:PLN02919  573 LAIDLLNNRLFISDSNHNRIV-------VTDLDGNFIVqIGSTGEEGLrdgSFEDA-----------TFNRPQGLAYNaK 634
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1045 NGLIYFVD--GTMIRKVD-QNGIISTLLGS----NDLTSARPLTcdtsmhiSQVrLEWPTDLAINPMDNSIYV------- 1110
Cdd:PLN02919  635 KNLLYVADteNHALREIDfVNETVRTLAGNgtkgSDYQGGKKGT-------SQV-LNSPWDVCFEPVNEKVYIamagqhq 706
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1111 ------LD-------------------------------------NNVVLQITENRQVR-----------IAAGRPMhcq 1136
Cdd:PLN02919  707 iweyniSDgvtrvfsgdgyernlngssgtstsfaqpsgislspdlKELYIADSESSSIRaldlktggsrlLAGGDPT--- 783
                         250       260       270       280       290       300       310       320
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1137 VPGVEYPVGKH---AVQTTLESATAIAVSYSGVLYITETDEKKINRIRQVTtdGEISLVAGIPsecdckndancdcyQSG 1213
Cdd:PLN02919  784 FSDNLFKFGDHdgvGSEVLLQHPLGVLCAKDGQIYVADSYNHKIKKLDPAT--KRVTTLAGTG--------------KAG 847
                         330       340       350       360
                  ....*....|....*....|....*....|....*....|..
gi 568954801 1214 --DGYAKDAKLNAPSSLAASPDGTLYIADLGNIRIRAVSKNK 1253
Cdd:PLN02919  848 fkDGKALKAQLSEPAGLALGENGRLFVADTNNSLIRYLDLNK 889
NHL cd05819
NHL repeat unit of beta-propeller proteins; The NHL(NCL-1, HT2A and LIN-41)-repeat is found in ...
1218-1344 8.75e-05

NHL repeat unit of beta-propeller proteins; The NHL(NCL-1, HT2A and LIN-41)-repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures. The repeats have a catalytic activity in Peptidyl-glycine alpha-amidating monooxygenase; proteolysis has shown that the Peptidyl-alpha-hydroxyglycine alpha-amidating lyase (PAL) activity is localized to the repeats. Tripartite motif-containing protein 32 interacts with the activation domain of Tat. This interaction is mediated by the NHL repeats.


Pssm-ID: 271320 [Multi-domain]  Cd Length: 269  Bit Score: 46.93  E-value: 8.75e-05
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1218 KDAKLNAPSSLAASPDGTLYIADLGNIRIRAVSKN-KPLLN-------SMNFYE---VASPTDQELYI----------FD 1276
Cdd:cd05819     3 GPGELNNPQGIAVDSSGNIYVADTGNNRIQVFDPDgNFITSfgsfgsgDGQFNEpagVAVDSDGNLYVadtgnhriqkFD 82
                          90       100       110       120       130       140       150
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....
gi 568954801 1277 INGTHQYTVSlVTGDYLYNFSY------SNDNDVtAVTDSNGNtlRIrrdpnrmpvRVVSPDNQVIwLTIGTNG 1344
Cdd:cd05819    83 PDGNFLASFG-GSGDGDGEFNGprgiavDSSGNI-YVADTGNH--RI---------QKFDPDGEFL-TTFGSGG 142
C_rich_MXAN6577 NF041328
MXAN_6577-like cysteine-rich domain;
409-511 1.28e-04

MXAN_6577-like cysteine-rich domain;


Pssm-ID: 469225 [Multi-domain]  Cd Length: 145  Bit Score: 44.36  E-value: 1.28e-04
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  409 CSVDCGSHGVCMGGSCRCEEGWT--GPAC--------NQRACHPRCAEHGTCKDGKCecsqgwngehctiahyldkivka 478
Cdd:NF041328   45 CGVACGAGQTCVAGACGCGPGTVacGGACvdtasdpaHCGACGAACAPGQVCEGGAC----------------------- 101
                          90       100       110       120
                  ....*....|....*....|....*....|....*....|
gi 568954801  479 dkigyKEGCP-GLCNSNGRCT-LDQNGWHC-----VCQPG 511
Cdd:NF041328  102 -----REACSeGLTRCGGACVdLATDPLHCgacgvACDPG 136
EGF_CA cd00054
Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular ...
341-370 2.86e-03

Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular (mostly animal) proteins. Many of these proteins require calcium for their biological function and calcium-binding sites have been found to be located at the N-terminus of particular EGF-like domains; calcium-binding may be crucial for numerous protein-protein interactions. Six conserved core cysteines form three disulfide bridges as in non calcium-binding EGF domains, whose structures are very similar. EGF_CA can be found in tandem repeat arrangements.


Pssm-ID: 238011  Cd Length: 38  Bit Score: 37.23  E-value: 2.86e-03
                          10        20        30
                  ....*....|....*....|....*....|....*
gi 568954801  341 DCLDPG-CSNHGVCIHGE----CHCNPGWGGSNCE 370
Cdd:cd00054     4 ECASGNpCQNGGTCVNTVgsyrCSCPPGYTGRNCE 38
EGF_CA cd00054
Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular ...
488-517 3.62e-03

Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular (mostly animal) proteins. Many of these proteins require calcium for their biological function and calcium-binding sites have been found to be located at the N-terminus of particular EGF-like domains; calcium-binding may be crucial for numerous protein-protein interactions. Six conserved core cysteines form three disulfide bridges as in non calcium-binding EGF domains, whose structures are very similar. EGF_CA can be found in tandem repeat arrangements.


Pssm-ID: 238011  Cd Length: 38  Bit Score: 37.23  E-value: 3.62e-03
                          10        20        30
                  ....*....|....*....|....*....|
gi 568954801  488 PGLCNSNGRCTLDQNGWHCVCQPGWRGAGC 517
Cdd:cd00054     8 GNPCQNGGTCVNTVGSYRCSCPPGYTGRNC 37
RHS_repeat pfam05593
RHS Repeat; RHS proteins contain extended repeat regions. These repeats often appear to be ...
1366-1397 7.04e-03

RHS Repeat; RHS proteins contain extended repeat regions. These repeats often appear to be involved in ligand binding. Note that this model may not find all the repeats in a protein and that it covers two RHS repeats. The 3D structure of an RHS-repeat-containing protein (the B and C components of an ABC toxin complex) has been determined. The RHS repeats form an extended strip of beta-sheet that spirals around to form a hollow shell, encapsulating the variable C-terminal domain.


Pssm-ID: 461685 [Multi-domain]  Cd Length: 37  Bit Score: 36.42  E-value: 7.04e-03
                           10        20        30
                   ....*....|....*....|....*....|..
gi 568954801  1366 GLLATKSDETGWTTFFDYDSEGRLTNVTFPTG 1397
Cdd:pfam05593    5 GRLTSVTDPDGRVTTYTYDAAGRLTAVTDPDG 36
EGF pfam00008
EGF-like domain; There is no clear separation between noise and signal. pfam00053 is very ...
491-514 9.32e-03

EGF-like domain; There is no clear separation between noise and signal. pfam00053 is very similar, but has 8 instead of 6 conserved cysteines. Includes some cytokine receptors. The EGF domain misses the N-terminus regions of the Ca2+ binding EGF domains (this is the main reason of discrepancy between swiss-prot domain start/end and Pfam). The family is hard to model due to many similar but different sub-types of EGF domains. Pfam certainly misses a number of EGF domains.


Pssm-ID: 394967  Cd Length: 31  Bit Score: 35.82  E-value: 9.32e-03
                           10        20
                   ....*....|....*....|....
gi 568954801   491 CNSNGRCTLDQNGWHCVCQPGWRG 514
Cdd:pfam00008    6 CSNGGTCVDTPGGYTCICPEGYTG 29
 
Name Accession Description Interval E-value
NHL_like_1 cd14953
Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat ...
890-1249 2.53e-41

Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat domains is found in a variety of domain architectures. The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.


Pssm-ID: 271323 [Multi-domain]  Cd Length: 323  Bit Score: 156.15  E-value: 2.53e-41
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  890 VSSIMGNGRRrsiscpscnGQADGNKLLA----PVALACGIDGSLYVGDF--NYVRRIFPSGNVTSVLELRNKDFRHSSN 963
Cdd:cd14953     1 VSTVAGSGTA---------GFSGGGGTAArfnsPSGVAVDAAGNLYVADRgnHRIRKITPDGVVTTVAGTGTAGFADGGG 71
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  964 PAHRYY----LATDPvTGDLYVSDTNTRRIYRpksLTGAKDLTknaeVVAGTGEqclpfdeARCGDGGKAVEATLMSPKG 1039
Cdd:cd14953    72 AAAQFNtpsgVAVDA-AGNLYVADTGNHRIRK---ITPDGVVS----TLAGTGT-------AGFSDDGGATAAQFNYPTG 136
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1040 MAIDKNGLIYFVDGT--MIRKVDQNGIISTLLGSNDLTSAR--PLTcdtsmhisQVRLEWPTDLAINPMDNsIYVLD--N 1113
Cdd:cd14953   137 VAVDAAGNLYVADTGnhRIRKITPDGVVTTVAGTGGAGYAGdgPAT--------AAQFNNPTGVAVDAAGN-LYVADrgN 207
                         250       260       270       280       290       300       310       320
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1114 NVVLQITENRQVRIAAGRPmhcqvpGVEYPVGKHAVQTTLESATAIAVSYSGVLYITETDEkkiNRIRQVTTDGEISLVA 1193
Cdd:cd14953   208 HRIRKITPDGVVTTVAGTG------TAGFSGDGGATAAQLNNPTGVAVDAAGNLYVADSGN---HRIRKITPAGVVTTVA 278
                         330       340       350       360       370
                  ....*....|....*....|....*....|....*....|....*....|....*..
gi 568954801 1194 GIPSEcdckndancdcyQSGD-GYAKDAKLNAPSSLAASPDGTLYIADLGNIRIRAV 1249
Cdd:cd14953   279 GGGAG------------FSGDgGPATSAQFNNPTGVAVDAAGNLYVADTGNNRIRKI 323
Tox-GHH pfam15636
GHH signature containing HNH/Endo VII superfamily nuclease toxin; A predicted toxin of the HNH ...
2366-2443 7.00e-40

GHH signature containing HNH/Endo VII superfamily nuclease toxin; A predicted toxin of the HNH/Endonuclease VII fold present in bacterial polymorphic toxin systems with a characteriztic sG[HQ]H signature motif. In bacterial polymorphic toxin systems, the toxin is exported by the type 2, type 6, type 7 or TcdB/TcaC-type secretion system. The metazoan teneurin proteins possess an inactive of this domain at their C-terminus.


Pssm-ID: 464783  Cd Length: 78  Bit Score: 142.75  E-value: 7.00e-40
                           10        20        30        40        50        60        70
                   ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*...
gi 568954801  2366 EEKARILEQARQRALARAWAREQQRVRDGEEGARLWTEGEKRQLLSAGKVQGYDGYYVLSVEQYPELADSANNIQFLR 2443
Cdd:pfam15636    1 EERKRLLEHAKKRAVREAWHRERQLLRNGLPGSRDWTDEEKEELLSTGSVPGYDGEYIHPVEQYPELADDPSNIRFRK 78
NHL_like_1 cd14953
Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat ...
970-1250 1.25e-38

Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat domains is found in a variety of domain architectures. The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.


Pssm-ID: 271323 [Multi-domain]  Cd Length: 323  Bit Score: 148.06  E-value: 1.25e-38
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  970 LATDPvTGDLYVSDTNTRRIYRpksltgakdLTKNAEV--VAGTGEqclpfdEARCGDGGKAveATLMSPKGMAIDKNGL 1047
Cdd:cd14953    28 VAVDA-AGNLYVADRGNHRIRK---------ITPDGVVttVAGTGT------AGFADGGGAA--AQFNTPSGVAVDAAGN 89
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1048 IYFVDGT--MIRKVDQNGIISTLLGsndlTSARPLTCDTSMhiSQVRLEWPTDLAINPMDNsIYVLD--NNVVLQITENR 1123
Cdd:cd14953    90 LYVADTGnhRIRKITPDGVVSTLAG----TGTAGFSDDGGA--TAAQFNYPTGVAVDAAGN-LYVADtgNHRIRKITPDG 162
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1124 QVRIAAGRPmhcqVPGveYPVGKHAVQTTLESATAIAVSYSGVLYITETDEkkiNRIRQVTTDGEISLVAGIPSEcdckn 1203
Cdd:cd14953   163 VVTTVAGTG----GAG--YAGDGPATAAQFNNPTGVAVDAAGNLYVADRGN---HRIRKITPDGVVTTVAGTGTA----- 228
                         250       260       270       280
                  ....*....|....*....|....*....|....*....|....*..
gi 568954801 1204 dancdcYQSGDGYAKDAKLNAPSSLAASPDGTLYIADLGNIRIRAVS 1250
Cdd:cd14953   229 ------GFSGDGGATAAQLNNPTGVAVDAAGNLYVADSGNHRIRKIT 269
RhsA COG3209
Uncharacterized conserved protein RhaS, contains 28 RHS repeats [General function prediction ...
1203-2143 4.86e-34

Uncharacterized conserved protein RhaS, contains 28 RHS repeats [General function prediction only];


Pssm-ID: 442442 [Multi-domain]  Cd Length: 1103  Bit Score: 144.13  E-value: 4.86e-34
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1203 NDANCDCYQSGDGYAKDAKLNAPSSLAASPDGTLYIADLGNIRIRAVSKNKPLLNSMNFYEVASPTDQELYIFDINGTHQ 1282
Cdd:COG3209   109 AAATASAGRLVSTGAGAGGTVTAATGGTLGATAGSATTGSTDGGRGGVAVTGLAGGGASAYGLTLGGAAAGPATGVGTGA 188
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1283 YTVSLVTGDYLYNFSYSNDNDVTAVTDSNGNTLRIRRDPNRMPVRVVSPDNQVIWLTIGTNGCLKSMTAQGLELVLFTYH 1362
Cdd:COG3209   189 VTLATGLAGSALLALGSGAILGGLAGAYSGSATTATGTALGTPASVAATVTGSATGAAGAGAAVATAATTLGGTTGAGTG 268
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1363 GNSGLLATK---SDETGWTTFFDYDSEGRLTNVTFPTGVVTNLHGDMDKAITVDIESSSREEDVSITSNLSSIDSFYTMV 1439
Cdd:COG3209   269 ASGAGLDAStgtGGAGGSNAAATAGGLGGAGLGSGGAGGGGTAGGTTTAAGTTGTAAVSGAADAGTTTTTGTGTGGTTTT 348
                         250       260       270       280       290       300       310       320
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1440 QDQLRNSYQIGYDGSLRIFYASGLDSHYQTEPHVLAGTANPTVAKRNMTLPGENGQNLVEWRFRKEQAQGKVNVFGRKLR 1519
Cdd:COG3209   349 VGGGGSLTLGGYGAAGGLTTSVGAGGGGSTSGSTTTVGGGGTATGSGGGSSTTGVGAGTTTTSTTGGDGGPATAAGALTA 428
                         330       340       350       360       370       380       390       400
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1520 VNGRNLLSVDFDRTTKTEKIYDDHRKFLLRIAYDTSGHPTLWLPSSKLMAVNVTYSSTGQIASIQRGTTSEKVDYDSQGR 1599
Cdd:COG3209   429 GGTATGTGTGGGGTTAGTDATTTTGGAGASGTLTTTGGAATGATTGGGTEAGTGGGTLTSGSAGATTLGTDTTLDDTLGG 508
                         410       420       430       440       450       460       470       480
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1600 IVSRVFADGKTWSYTYLEKSMVLLLHSQRQYIFEYDMWDRLSAITMPSVARHTMQTIRSIGYYRNIYNPPESNASIITDY 1679
Cdd:COG3209   509 TTTTTAGARGLVVTTGTTLTLGTTTTATLSATDATGTGDTTTTGTVGTGTSTGTGGTGTVTTTGDGTGGASTTTGTTGGT 588
                         490       500       510       520       530       540       550       560
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1680 NEEGLLLQTAFLGTSRRVLFKYRRQTRLSEILYDSTRVSFTYDETAGVLKTVNLQSDGFicTIRYRQIGPLIDRQIFRFS 1759
Cdd:COG3209   589 ATTTTVTTTTTTSTAGTTTTTTSGYTRAGLTLTLGTGTASGLERATASTGSTTGGTTGT--GVTTTGTTTTRATGTTGTG 666
                         570       580       590       600       610       620       630       640
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1760 EDGMVNARFDYSYDNSFRVTSMQGVINETPLPIDLYQFDDISGKVEQFGKFGVIYYDINQIISTAVMTYTKHFDAHGRIK 1839
Cdd:COG3209   667 TGVTAGLTTLATGGTTVGGGTGTTSTATTGATTGGTETGTTVTTLAGGTTTRLGTTTTGGGGGTTTDGTGTGGTTGTLTT 746
                         650       660       670       680       690       700       710       720
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1840 EIQYEifRSLMYWITIQYDNMGRVTKREIKIGPFANTTKYAYEYDVDGQLQTVYLNEKIMWRYNYDLNGNLHLLNPSSSA 1919
Cdd:COG3209   747 TSTTT--TTTAGALTYTYDALGRLTSETTPGGVTQGTYTTRYTYDALGRLTSVTYPDGETVTYTYDALGRLTSVITVGSG 824
                         730       740       750       760       770       780       790       800
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1920 RLTPL-----RYDLRDRITRLgdvqyrldEDGFLRQRGTEIFEYSSKGLLTRVYSKGSGWTviYRYDGLGRRVSSKTSLG 1994
Cdd:COG3209   825 GGTDLqdrtyTYDAAGNITSI--------TDALRAGTLTQTYTYDALGRLTSATDPGTTES--YTYDANGNLTSRTDGGT 894
                         810       820       830       840       850       860       870       880
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1995 QHLQFFYADLtyPTRITHvynhSSSEITSLYYDLQGHlfameissgdefyiaSDNTGTPLAVFSSNGLMLKQIQYTAYGE 2074
Cdd:COG3209   895 TTYTYDALGR--LVSVTK----PDGTTTTYTYDALGH---------------TDHLGSVRALTDASGQVVWRYDYDPFGN 953
                         890       900       910       920       930       940
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....
gi 568954801 2075 IYFDSNVDFQLVIGFHGGLYDPLTKLIHFGERDYDILAGRWTTPDieiwkRIGKDPAPfNLYMFRNNNP 2143
Cdd:COG3209   954 LLAETSGAAANPLRFTGQEYDAETGLYYNGARYYDPALGRFLSPD-----PIGLAGGL-NLYAYVGNNP 1016
NHL_like_1 cd14953
Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat ...
1007-1250 5.48e-32

Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat domains is found in a variety of domain architectures. The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.


Pssm-ID: 271323 [Multi-domain]  Cd Length: 323  Bit Score: 128.80  E-value: 5.48e-32
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1007 VVAGTGeqclpfdeARCGDGGKAVEATLMSPKGMAIDKNGLIYFVDGT--MIRKVDQNGIISTLLG------SNDLTSAr 1078
Cdd:cd14953     3 TVAGSG--------TAGFSGGGGTAARFNSPSGVAVDAAGNLYVADRGnhRIRKITPDGVVTTVAGtgtagfADGGGAA- 73
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1079 pltcdtsmhisqVRLEWPTDLAINPMDNsIYVLD--NNVVLQITENRQVRIAAGrpmhcqVPGVEYPVGKHAVQTTLESA 1156
Cdd:cd14953    74 ------------AQFNTPSGVAVDAAGN-LYVADtgNHRIRKITPDGVVSTLAG------TGTAGFSDDGGATAAQFNYP 134
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1157 TAIAVSYSGVLYITETDEkkiNRIRQVTTDGEISLVAGIPSEcdckndancdcYQSGDGYAKDAKLNAPSSLAASPDGTL 1236
Cdd:cd14953   135 TGVAVDAAGNLYVADTGN---HRIRKITPDGVVTTVAGTGGA-----------GYAGDGPATAAQFNNPTGVAVDAAGNL 200
                         250
                  ....*....|....
gi 568954801 1237 YIADLGNIRIRAVS 1250
Cdd:cd14953   201 YVADRGNHRIRKIT 214
NHL cd05819
NHL repeat unit of beta-propeller proteins; The NHL(NCL-1, HT2A and LIN-41)-repeat is found in ...
909-1247 4.46e-19

NHL repeat unit of beta-propeller proteins; The NHL(NCL-1, HT2A and LIN-41)-repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures. The repeats have a catalytic activity in Peptidyl-glycine alpha-amidating monooxygenase; proteolysis has shown that the Peptidyl-alpha-hydroxyglycine alpha-amidating lyase (PAL) activity is localized to the repeats. Tripartite motif-containing protein 32 interacts with the activation domain of Tat. This interaction is mediated by the NHL repeats.


Pssm-ID: 271320 [Multi-domain]  Cd Length: 269  Bit Score: 89.69  E-value: 4.46e-19
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  909 GQADGnKLLAPVALACGIDGSLYVGDF--NYVRRIFPSGNVTSVLELRNKDFRHSSNPAHryyLATDPvTGDLYVSDTNT 986
Cdd:cd05819     1 GTGPG-ELNNPQGIAVDSSGNIYVADTgnNRIQVFDPDGNFITSFGSFGSGDGQFNEPAG---VAVDS-DGNLYVADTGN 75
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  987 RRIYRpksltgakdLTKNAEVVAGTGeqclpfdearcGDGGKAVEatLMSPKGMAIDKNGLIYFVDgTM---IRKVDQNG 1063
Cdd:cd05819    76 HRIQK---------FDPDGNFLASFG-----------GSGDGDGE--FNGPRGIAVDSSGNIYVAD-TGnhrIQKFDPDG 132
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1064 IISTLLGSNDLTSARpltcdtsmhisqvrLEWPTDLAINPmDNSIYVLDnnvvlqiTENRQVRI--AAGRPMHcQVPGVE 1141
Cdd:cd05819   133 EFLTTFGSGGSGPGQ--------------FNGPTGVAVDS-DGNIYVAD-------TGNHRIQVfdPDGNFLT-TFGSTG 189
                         250       260       270       280       290       300       310       320
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1142 YPVGKhavqttLESATAIAVSYSGVLYITETDEkkiNRIRQVTTDGEISlvagipsecdckndancdcYQSGDGYAKDAK 1221
Cdd:cd05819   190 TGPGQ------FNYPTGIAVDSDGNIYVADSGN---NRVQVFDPDGAGF-------------------GGNGNFLGSDGQ 241
                         330       340
                  ....*....|....*....|....*.
gi 568954801 1222 LNAPSSLAASPDGTLYIADLGNIRIR 1247
Cdd:cd05819   242 FNRPSGLAVDSDGNLYVADTGNNRIQ 267
Ten_N pfam06484
Teneurin Intracellular Region; This family is found in the intracellular N-terminal region of ...
1-36 1.13e-17

Teneurin Intracellular Region; This family is found in the intracellular N-terminal region of the Teneurin family of proteins. These proteins are 'pair-rule' genes and are involved in tissue patterning, specifically probably neural patterning. The intracellular domain is cleaved in response to homophilic interaction of the extracellular domain, and translocates to the nucleus. Here it probably carries out to some transcriptional regulatory activity. The length of this region and the conservation suggests that there may be two structural domains here (personal obs:C Yeats).


Pssm-ID: 461932 [Multi-domain]  Cd Length: 367  Bit Score: 87.34  E-value: 1.13e-17
                           10        20        30
                   ....*....|....*....|....*....|....*.
gi 568954801     1 MASGSVYSPPTRPLPRNTLSRSAFKFKKSSKYCSWR 36
Cdd:pfam06484  332 LTSGTVYSPPPRPLPRNTFSRPAFKLKKPYKYCSWK 367
NHL cd05819
NHL repeat unit of beta-propeller proteins; The NHL(NCL-1, HT2A and LIN-41)-repeat is found in ...
908-1180 1.04e-15

NHL repeat unit of beta-propeller proteins; The NHL(NCL-1, HT2A and LIN-41)-repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures. The repeats have a catalytic activity in Peptidyl-glycine alpha-amidating monooxygenase; proteolysis has shown that the Peptidyl-alpha-hydroxyglycine alpha-amidating lyase (PAL) activity is localized to the repeats. Tripartite motif-containing protein 32 interacts with the activation domain of Tat. This interaction is mediated by the NHL repeats.


Pssm-ID: 271320 [Multi-domain]  Cd Length: 269  Bit Score: 79.67  E-value: 1.04e-15
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  908 NGQADGNkLLAPVALACGIDGSLYVGDF--NYVRRIFPSGNVTSVLELRNKDFRHSSNPahrYYLATDPvTGDLYVSDTN 985
Cdd:cd05819    47 FGSGDGQ-FNEPAGVAVDSDGNLYVADTgnHRIQKFDPDGNFLASFGGSGDGDGEFNGP---RGIAVDS-SGNIYVADTG 121
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  986 TRRIYRpksltgakdLTKNAEVVAGTGeqclpfdearcgdGGKAVEATLMSPKGMAIDKNGLIYFVDGT--MIRKVDQNG 1063
Cdd:cd05819   122 NHRIQK---------FDPDGEFLTTFG-------------SGGSGPGQFNGPTGVAVDSDGNIYVADTGnhRIQVFDPDG 179
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1064 IISTLLGSNDLTSARpltcdtsmhisqvrLEWPTDLAINPMDNsIYVLD--NNVVLQITENRQVRIAAGRPMhCQVPGVE 1141
Cdd:cd05819   180 NFLTTFGSTGTGPGQ--------------FNYPTGIAVDSDGN-IYVADsgNNRVQVFDPDGAGFGGNGNFL-GSDGQFN 243
                         250       260       270
                  ....*....|....*....|....*....|....*....
gi 568954801 1142 YPVGkhavqttlesataIAVSYSGVLYITETDEKKINRI 1180
Cdd:cd05819   244 RPSG-------------LAVDSDGNLYVADTGNNRIQVF 269
NHL_PKND_like cd14952
NHL repeat domain of the protein kinase PknD; PknD is a mycobacterial transmembrane protein ...
970-1246 3.31e-09

NHL repeat domain of the protein kinase PknD; PknD is a mycobacterial transmembrane protein with a cytosolic kinase domain and an extracellular sensor domain that contains NHL repeats. It plays a key role in the development of central nervous system tuberculosis, by mediating the invasion of host brain endothelia. The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.


Pssm-ID: 271322 [Multi-domain]  Cd Length: 247  Bit Score: 59.91  E-value: 3.31e-09
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  970 LATDPvTGDLYVSDTNTRRIYRpksltgakdltknaeVVAGTGEQ-CLPFDEarcgdggkaveatLMSPKGMAIDKNGLI 1048
Cdd:cd14952    15 VAVDA-AGNVYVADSGNNRVLK---------------LAAGSTTQtVLPFTG-------------LYQPQGVAVDAAGTV 65
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1049 YFVDGtmirkvDQNGIISTLLGSNDLTsARPLTcdtsmhisqvRLEWPTDLAINPMDNsIYVLDNnvvlqiTENRQVRIA 1128
Cdd:cd14952    66 YVTDF------GNNRVLKLAAGSTTQT-VLPFT----------GLNDPTGVAVDAAGN-VYVADT------GNNRVLKLA 121
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1129 AGRPMHCQVPgveypvgkhavQTTLESATAIAVSYSGVLYITETDEkkiNRIRQvttdgeisLVAGipsecdckndANCD 1208
Cdd:cd14952   122 AGSNTQTVLP-----------FTGLSNPDGVAVDGAGNVYVTDTGN---NRVLK--------LAAG----------STTQ 169
                         250       260       270
                  ....*....|....*....|....*....|....*...
gi 568954801 1209 CYQSGDGyakdakLNAPSSLAASPDGTLYIADLGNIRI 1246
Cdd:cd14952   170 TVLPFTG------LNSPSGVAVDTAGNVYVTDHGNNRV 201
Rhs_assc_core TIGR03696
RHS repeat-associated core domain; This model represents a conserved unique core sequence ...
2069-2143 3.40e-09

RHS repeat-associated core domain; This model represents a conserved unique core sequence shared by large numbers of proteins. It is occasional in the Archaea Methanosarcina barkeri) but common in bacteria and eukaryotes. Most fall into two large classes. One class consists of long proteins in which two classes of repeats are abundant: an FG-GAP repeat (pfam01839) class, and an RHS repeat (pfam05593) or YD repeat (TIGR01643). This class includes secreted bacterial insecticidal toxins and intercellular signalling proteins such as the teneurins in animals. The other class consists of uncharacterized proteins shorter than 400 amino acids, where this core domain of about 75 amino acids tends to occur in the N-terminal half. Over twenty such proteins are found in Pseudomonas putida alone; little sequence similarity or repeat structure is found among these proteins outside the region modeled by this domain.


Pssm-ID: 274730 [Multi-domain]  Cd Length: 77  Bit Score: 55.20  E-value: 3.40e-09
                           10        20        30        40        50        60        70
                   ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*
gi 568954801  2069 YTAYGEIYFDSNVDFQLvIGFHGGLYDPLTKLIHFGERDYDILAGRWTTPDieiwkRIGKDpAPFNLYMFRNNNP 2143
Cdd:TIGR03696    1 YDPYGEVLSESGAAPNP-LRFTGQYYDAETGLYYNGARYYDPELGRFLSPD-----PIGLG-GGLNLYAYVGNNP 68
Vgb COG4257
Streptogramin lyase [Defense mechanisms];
919-1197 3.43e-09

Streptogramin lyase [Defense mechanisms];


Pssm-ID: 443399 [Multi-domain]  Cd Length: 270  Bit Score: 60.42  E-value: 3.43e-09
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  919 PVALACGIDGSLYVGDF--NYVRRIFP-SGNVTsvlelrnkdfRHSSNPAHRYY-LATDPvTGDLYVSDTNTRRIYRpks 994
Cdd:COG4257    19 PRDVAVDPDGAVWFTDQggGRIGRLDPaTGEFT----------EYPLGGGSGPHgIAVDP-DGNLWFTDNGNNRIGR--- 84
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  995 LTGAkdlTKNAEVVAGTGEQCLPFdearcgdggkaveatlmspkGMAIDKNGLIYFVDGT--MIRKVD-QNGIISTLLGS 1071
Cdd:COG4257    85 IDPK---TGEITTFALPGGGSNPH--------------------GIAFDPDGNLWFTDQGgnRIGRLDpATGEVTEFPLP 141
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1072 NDLTSARPLTCD--------------------TSMHISQVRLE----WPTDLAINPmDNSIYVLD--NNVVLQITEnrqv 1125
Cdd:COG4257   142 TGGAGPYGIAVDpdgnlwvtdfganaigridpDTGTLTEYALPtpgaGPRGLAVDP-DGNLWVADtgSGRIGRFDP---- 216
                         250       260       270       280       290       300       310
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|..
gi 568954801 1126 riAAGRpmhcqvpgveypVGKHAVQTTLESATAIAVSYSGVLYITETDekkINRIRQVTTDGEISLVAgIPS 1197
Cdd:COG4257   217 --KTGT------------VTEYPLPGGGARPYGVAVDGDGRVWFAESG---ANRIVRFDPDTELTEYV-LPS 270
DUF5885 pfam19232
Family of unknown function (DUF5885); This is a family of uncharacterized proteins of unknown ...
250-435 1.99e-08

Family of unknown function (DUF5885); This is a family of uncharacterized proteins of unknown function found in viruses.


Pssm-ID: 437064  Cd Length: 265  Bit Score: 57.71  E-value: 1.99e-08
                           10        20        30        40        50        60        70        80
                   ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801   250 CHGNGECvsGTCH-CFPGFLGPDCSraACPVLCSGNGQ----------YSKGRC----LCFSGwkgTECDVPTTQCI-DP 313
Cdd:pfam19232   34 CTTDAQC--GTCMtCVAGACTPKAS--CCGGVTCGAGQtcdaktntcvYVKGYCsadhPCPSG---SACDTAKNACIaQP 106
                           90       100       110       120       130       140       150       160
                   ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801   314 QCG---GRGiCIMG-------------------SCACNSGYK-GENCE--------EADCLDP---------------GC 347
Cdd:pfam19232  107 PYGpdsGKG-CVRGfgawiweldpatnsgvwrcRCANGSLYNsAHECSpladqtlcAAENLDPnalvpassvpafaayGW 185
                          170       180       190       200       210       220       230       240
                   ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801   348 SNHGVCIH-------------GECHCNPGWGGSNCEILKTmcadqCSGHGTYLQESGSCTCDPNWTGpdcsneicsvdcg 414
Cdd:pfam19232  186 GNQPVLINkstagaavpsplaGVCPCKPGWAGGSCTEDRT-----CNGRGTWNETTGQCACNIDFSG------------- 247
                          250       260
                   ....*....|....*....|....
gi 568954801   415 shgvcmGGSCRCEEG---WTGPAC 435
Cdd:pfam19232  248 ------HNSCGDDNNctsWTGPRC 265
NHL_like_2 cd14957
Uncharacterized NHL-repeat domain in bacterial and archaeal proteins; The NHL (NCL-1, HT2A and ...
1034-1344 3.58e-08

Uncharacterized NHL-repeat domain in bacterial and archaeal proteins; The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.


Pssm-ID: 271327 [Multi-domain]  Cd Length: 280  Bit Score: 57.28  E-value: 3.58e-08
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1034 LMSPKGMAIDKNGLIYFVD--GTMIRKVDQNGIISTLLGSNDltsarpltcdtsmhISQVRLEWPTDLAINPMDNsIYVL 1111
Cdd:cd14957    17 FNTPRGIAVDSAGNIYVADtgNNRIQVFTSSGVYSYSIGSGG--------------TGSGQFNSPYGIAVDSNGN-IYVA 81
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1112 DNNvvlqitENR-QVRIAAGrpmhcqvpGVEYPVGKHAVQTT-LESATAIAVSYSGVLYITETDEkkiNRIRQVTTDGEI 1189
Cdd:cd14957    82 DTD------NNRiQVFNSSG--------VYQYSIGTGGSGDGqFNGPYGIAVDSNGNIYVADTGN---HRIQVFTSSGTF 144
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1190 slvagipsecdckndancdCYQSGDGYAKDAKLNAPSSLAASPDGTLYIADLGNIRIRavsknkpllnsmnfyevasptd 1269
Cdd:cd14957   145 -------------------SYSIGSGGTGPGQFNGPQGIAVDSDGNIYVADTGNHRIQ---------------------- 183
                         250       260       270       280       290       300       310
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*.
gi 568954801 1270 qelyIFDINGTHQYTV-SLVTGDYLynFSYSNDNDVtavtDSNGNTLRIRRDPNRmpVRVVSPDNqVIWLTIGTNG 1344
Cdd:cd14957   184 ----VFTSSGTFQYTFgSSGSGPGQ--FSDPYGIAV----DSDGNIYVADTGNHR--IQVFTSSG-AYQYSIGTSG 246
acid_disulf_rpt NF033662
acidic double-disulfide repeat; The acidic double-disulfide repeat is an Asp-rich repeat with ...
520-550 9.90e-08

acidic double-disulfide repeat; The acidic double-disulfide repeat is an Asp-rich repeat with four nearly invariant Cys residues in a repeat length of about 35 amino acids.


Pssm-ID: 411265 [Multi-domain]  Cd Length: 32  Bit Score: 49.82  E-value: 9.90e-08
                          10        20        30
                  ....*....|....*....|....*....|.
gi 568954801  520 AMETLCTDSKDNEGDGLIDCMDPDCCLQSSC 550
Cdd:NF033662    2 ATDTTCSDGIDNDGDGLTDCADPDCAGNPVC 32
C_rich_MXAN6577 NF041328
MXAN_6577-like cysteine-rich domain;
301-455 6.92e-07

MXAN_6577-like cysteine-rich domain;


Pssm-ID: 469225 [Multi-domain]  Cd Length: 145  Bit Score: 50.91  E-value: 6.92e-07
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  301 TECDVPTTQCIDPQ--CGGRGICIMGScACNSGykgeNCEEAdcldpgCSNHGVCIHGECHCNPGwggsnceilKTMCAD 378
Cdd:NF041328   12 AGCPEPGAVCPEGLsvCGGACVDLRSD-PSNCG----ACGVA------CGAGQTCVAGACGCGPG---------TVACGG 71
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  379 QCSGHGTylqesgsctcDPNWTGPdcsneiCSVDCGSHGVCMGGSCR--CEEGWT--GPAC--------NQRACHPRCAE 446
Cdd:NF041328   72 ACVDTAS----------DPAHCGA------CGAACAPGQVCEGGACReaCSEGLTrcGGACvdlatdplHCGACGVACDP 135

                  ....*....
gi 568954801  447 HGTCKDGKC 455
Cdd:NF041328  136 GESCRGGAC 144
NHL_like_3 cd14956
Uncharacterized NHL-repeat domain in bacterial proteins; The NHL (NCL-1, HT2A and LIN-41) ...
977-1251 1.14e-06

Uncharacterized NHL-repeat domain in bacterial proteins; The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.


Pssm-ID: 271326 [Multi-domain]  Cd Length: 274  Bit Score: 52.67  E-value: 1.14e-06
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  977 GDLYVSDTNTRRIyrpksltgakdltknaEVVAGTGEQCLPFDEARCGDGGkaveatLMSPKGMAIDKNGLIYFVDGT-- 1054
Cdd:cd14956    24 DNVYVADARNGRI----------------QVFDKDGTFLRRFGTTGDGPGQ------FGRPRGLAVDKDGWLYVADYWgd 81
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1055 MIRKVDQNGIISTLLGSNdltSARPLTCDTsmhisqvrlewPTDLAINPmDNSIYVLD--NNVVLQITENRQVRIAAGRP 1132
Cdd:cd14956    82 RIQVFTLTGELQTIGGSS---GSGPGQFNA-----------PRGVAVDA-DGNLYVADfgNQRIQKFDPDGSFLRQWGGT 146
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1133 mhcqvpgvEYPVGKhavqttLESATAIAVSYSGVLYITETdekKINRIRQVTTDGEISLVAGIPSecdckndancdcyqS 1212
Cdd:cd14956   147 --------GIEPGS------FNYPRGVAVDPDGTLYVADT---YNDRIQVFDNDGAFLRKWGGRG--------------T 195
                         250       260       270
                  ....*....|....*....|....*....|....*....
gi 568954801 1213 GDGyakdaKLNAPSSLAASPDGTLYIADLGNIRIRAVSK 1251
Cdd:cd14956   196 GPG-----QFNYPYGIAIDPDGNVFVADFGNNRIQKFTA 229
Vgb COG4257
Streptogramin lyase [Defense mechanisms];
968-1247 1.73e-06

Streptogramin lyase [Defense mechanisms];


Pssm-ID: 443399 [Multi-domain]  Cd Length: 270  Bit Score: 51.94  E-value: 1.73e-06
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  968 YYLATDPvTGDLYVSDTNTRRIYRpksltgakdltknaeVVAGTGEqclpFDEARCGDGGkaveatlmSPKGMAIDKNGL 1047
Cdd:COG4257    20 RDVAVDP-DGAVWFTDQGGGRIGR---------------LDPATGE----FTEYPLGGGS--------GPHGIAVDPDGN 71
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1048 IYFVDGT--MIRKVD-QNGIISTLLGSNDLTSarpltcdtsmhisqvrlewPTDLAINPmDNSIYVLD--NNVVLQIT-E 1121
Cdd:COG4257    72 LWFTDNGnnRIGRIDpKTGEITTFALPGGGSN-------------------PHGIAFDP-DGNLWFTDqgGNRIGRLDpA 131
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1122 NRQVRiaagrpmhcqvpgvEYPVGKHAVQTTlesatAIAVSYSGVLYITETdekKINRIRQVTTD-GEISLvagipsecd 1200
Cdd:COG4257   132 TGEVT--------------EFPLPTGGAGPY-----GIAVDPDGNLWVTDF---GANAIGRIDPDtGTLTE--------- 180
                         250       260       270       280
                  ....*....|....*....|....*....|....*....|....*..
gi 568954801 1201 ckndancdcyqsgdgYAKDAKLNAPSSLAASPDGTLYIADLGNIRIR 1247
Cdd:COG4257   181 ---------------YALPTPGAGPRGLAVDPDGNLWVADTGSGRIG 212
NHL_PKND_like cd14952
NHL repeat domain of the protein kinase PknD; PknD is a mycobacterial transmembrane protein ...
916-1177 3.63e-06

NHL repeat domain of the protein kinase PknD; PknD is a mycobacterial transmembrane protein with a cytosolic kinase domain and an extracellular sensor domain that contains NHL repeats. It plays a key role in the development of central nervous system tuberculosis, by mediating the invasion of host brain endothelia. The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.


Pssm-ID: 271322 [Multi-domain]  Cd Length: 247  Bit Score: 50.67  E-value: 3.63e-06
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  916 LLAPVALACGIDGSLYVGDF--NYVRRIFPSGNVTSVLElrnkdFRHSSNPAHryyLATDPVtGDLYVSDTNTRRIyrpk 993
Cdd:cd14952    51 LYQPQGVAVDAAGTVYVTDFgnNRVLKLAAGSTTQTVLP-----FTGLNDPTG---VAVDAA-GNVYVADTGNNRV---- 117
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  994 sltgakdltknAEVVAGTGEQC-LPFdearcgdggkaveATLMSPKGMAIDKNGLIYFVDGtmirkvDQNGIISTLLGSN 1072
Cdd:cd14952   118 -----------LKLAAGSNTQTvLPF-------------TGLSNPDGVAVDGAGNVYVTDT------GNNRVLKLAAGST 167
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1073 DLTsARPLTCDTSmhisqvrlewPTDLAINPMDNsIYVLDNNvvlqitENRQVRIAAGRPMHCQVP--GVEYPVGkhavq 1150
Cdd:cd14952   168 TQT-VLPFTGLNS----------PSGVAVDTAGN-VYVTDHG------NNRVLKLAAGSTTPTVLPftGLNGPLG----- 224
                         250       260
                  ....*....|....*....|....*..
gi 568954801 1151 ttlesataIAVSYSGVLYITETDEKKI 1177
Cdd:cd14952   225 --------VAVDAAGNVYVADRGNDRV 243
DSL pfam01414
Delta serrate ligand; This family has been redefined to correspond to the EGF-like domain ...
426-469 1.61e-05

Delta serrate ligand; This family has been redefined to correspond to the EGF-like domain defined by structure.


Pssm-ID: 460202  Cd Length: 46  Bit Score: 44.15  E-value: 1.61e-05
                           10        20        30        40
                   ....*....|....*....|....*....|....*....|....*..
gi 568954801   426 CEEGWTGPACNqRACHPRCAE--HGTC-KDGKCECSQGWNGEHCTIA 469
Cdd:pfam01414    1 CDENYYGSTCS-KFCRPRDDKfgHYTCdANGNKVCLPGWTGPYCDKP 46
PLN02919 PLN02919
haloacid dehalogenase-like hydrolase family protein
970-1253 3.45e-05

haloacid dehalogenase-like hydrolase family protein


Pssm-ID: 215497 [Multi-domain]  Cd Length: 1057  Bit Score: 49.46  E-value: 3.45e-05
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  970 LATDPVTGDLYVSDTNTRRIYrpksltgAKDLTKNAEV-VAGTGEQCL---PFDEArcgdggkaveaTLMSPKGMAID-K 1044
Cdd:PLN02919  573 LAIDLLNNRLFISDSNHNRIV-------VTDLDGNFIVqIGSTGEEGLrdgSFEDA-----------TFNRPQGLAYNaK 634
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1045 NGLIYFVD--GTMIRKVD-QNGIISTLLGS----NDLTSARPLTcdtsmhiSQVrLEWPTDLAINPMDNSIYV------- 1110
Cdd:PLN02919  635 KNLLYVADteNHALREIDfVNETVRTLAGNgtkgSDYQGGKKGT-------SQV-LNSPWDVCFEPVNEKVYIamagqhq 706
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1111 ------LD-------------------------------------NNVVLQITENRQVR-----------IAAGRPMhcq 1136
Cdd:PLN02919  707 iweyniSDgvtrvfsgdgyernlngssgtstsfaqpsgislspdlKELYIADSESSSIRaldlktggsrlLAGGDPT--- 783
                         250       260       270       280       290       300       310       320
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1137 VPGVEYPVGKH---AVQTTLESATAIAVSYSGVLYITETDEKKINRIRQVTtdGEISLVAGIPsecdckndancdcyQSG 1213
Cdd:PLN02919  784 FSDNLFKFGDHdgvGSEVLLQHPLGVLCAKDGQIYVADSYNHKIKKLDPAT--KRVTTLAGTG--------------KAG 847
                         330       340       350       360
                  ....*....|....*....|....*....|....*....|..
gi 568954801 1214 --DGYAKDAKLNAPSSLAASPDGTLYIADLGNIRIRAVSKNK 1253
Cdd:PLN02919  848 fkDGKALKAQLSEPAGLALGENGRLFVADTNNSLIRYLDLNK 889
NHL_like_1 cd14953
Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat ...
1212-1252 7.32e-05

Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat domains is found in a variety of domain architectures. The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.


Pssm-ID: 271323 [Multi-domain]  Cd Length: 323  Bit Score: 47.52  E-value: 7.32e-05
                          10        20        30        40
                  ....*....|....*....|....*....|....*....|.
gi 568954801 1212 SGDGYAKDAKLNAPSSLAASPDGTLYIADLGNIRIRAVSKN 1252
Cdd:cd14953    12 FSGGGGTAARFNSPSGVAVDAAGNLYVADRGNHRIRKITPD 52
NHL cd05819
NHL repeat unit of beta-propeller proteins; The NHL(NCL-1, HT2A and LIN-41)-repeat is found in ...
1218-1344 8.75e-05

NHL repeat unit of beta-propeller proteins; The NHL(NCL-1, HT2A and LIN-41)-repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures. The repeats have a catalytic activity in Peptidyl-glycine alpha-amidating monooxygenase; proteolysis has shown that the Peptidyl-alpha-hydroxyglycine alpha-amidating lyase (PAL) activity is localized to the repeats. Tripartite motif-containing protein 32 interacts with the activation domain of Tat. This interaction is mediated by the NHL repeats.


Pssm-ID: 271320 [Multi-domain]  Cd Length: 269  Bit Score: 46.93  E-value: 8.75e-05
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1218 KDAKLNAPSSLAASPDGTLYIADLGNIRIRAVSKN-KPLLN-------SMNFYE---VASPTDQELYI----------FD 1276
Cdd:cd05819     3 GPGELNNPQGIAVDSSGNIYVADTGNNRIQVFDPDgNFITSfgsfgsgDGQFNEpagVAVDSDGNLYVadtgnhriqkFD 82
                          90       100       110       120       130       140       150
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....
gi 568954801 1277 INGTHQYTVSlVTGDYLYNFSY------SNDNDVtAVTDSNGNtlRIrrdpnrmpvRVVSPDNQVIwLTIGTNG 1344
Cdd:cd05819    83 PDGNFLASFG-GSGDGDGEFNGprgiavDSSGNI-YVADTGNH--RI---------QKFDPDGEFL-TTFGSGG 142
DSL pfam01414
Delta serrate ligand; This family has been redefined to correspond to the EGF-like domain ...
395-437 9.52e-05

Delta serrate ligand; This family has been redefined to correspond to the EGF-like domain defined by structure.


Pssm-ID: 460202  Cd Length: 46  Bit Score: 41.84  E-value: 9.52e-05
                           10        20        30        40
                   ....*....|....*....|....*....|....*....|....*.
gi 568954801   395 CDPNWTGPDCSNEiCSV--DCGSHGVC-MGGSCRCEEGWTGPACNQ 437
Cdd:pfam01414    1 CDENYYGSTCSKF-CRPrdDKFGHYTCdANGNKVCLPGWTGPYCDK 45
C_rich_MXAN6577 NF041328
MXAN_6577-like cysteine-rich domain;
409-511 1.28e-04

MXAN_6577-like cysteine-rich domain;


Pssm-ID: 469225 [Multi-domain]  Cd Length: 145  Bit Score: 44.36  E-value: 1.28e-04
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  409 CSVDCGSHGVCMGGSCRCEEGWT--GPAC--------NQRACHPRCAEHGTCKDGKCecsqgwngehctiahyldkivka 478
Cdd:NF041328   45 CGVACGAGQTCVAGACGCGPGTVacGGACvdtasdpaHCGACGAACAPGQVCEGGAC----------------------- 101
                          90       100       110       120
                  ....*....|....*....|....*....|....*....|
gi 568954801  479 dkigyKEGCP-GLCNSNGRCT-LDQNGWHC-----VCQPG 511
Cdd:NF041328  102 -----REACSeGLTRCGGACVdLATDPLHCgacgvACDPG 136
NHL_like_5 cd14963
Uncharacterized NHL-repeat domain in bacterial proteins; The NHL (NCL-1, HT2A and LIN-41) ...
915-1172 2.76e-04

Uncharacterized NHL-repeat domain in bacterial proteins; The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.


Pssm-ID: 271333 [Multi-domain]  Cd Length: 268  Bit Score: 45.36  E-value: 2.76e-04
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  915 KLLAPVALACGIDGSLYVGDFnYVRRI--F-PSGNVTSVLElRNKDFRHSSNPAHryyLATDpvTGDLYVSDTNTRRIYr 991
Cdd:cd14963    54 EFKYPYGIAVDSDGNIYVADL-YNGRIqvFdPDGKFLKYFP-EKKDRVKLISPAG---LAID--DGKLYVSDVKKHKVI- 125
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  992 pksltgakdltknaeVVAGTGEQCLPFdearcGDGGKAvEATLMSPKGMAIDKNGLIYFVD--GTMIRKVDQNG-IISTL 1068
Cdd:cd14963   126 ---------------VFDLEGKLLLEF-----GKPGSE-PGELSYPNGIAVDEDGNIYVADsgNGRIQVFDKNGkFIKEL 184
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1069 LGSNDLTSArpltcdtsmhisqvrLEWPTDLAINPmDNSIYVLDN--NVVLQITENRQVRIAAGRpmhcqvPGVEypvgk 1146
Cdd:cd14963   185 NGSPDGKSG---------------FVNPRGIAVDP-DGNLYVVDNlsHRVYVFDEQGKELFTFGG------RGKD----- 237
                         250       260
                  ....*....|....*....|....*.
gi 568954801 1147 havQTTLESATAIAVSYSGVLYITET 1172
Cdd:cd14963   238 ---DGQFNLPNGLFIDDDGRLYVTDR 260
EGF_2 pfam07974
EGF-like domain; This family contains EGF domains found in a variety of extracellular proteins.
250-272 5.56e-04

EGF-like domain; This family contains EGF domains found in a variety of extracellular proteins.


Pssm-ID: 400365  Cd Length: 26  Bit Score: 39.25  E-value: 5.56e-04
                           10        20
                   ....*....|....*....|....*
gi 568954801   250 CHGNGECVS--GTCHCFPGFLGPDC 272
Cdd:pfam07974    2 CSGRGTCVNqcGKCVCDSGYQGATC 26
Keratin_B2 pfam01500
Keratin, high sulfur B2 protein; High sulfur proteins are cysteine-rich proteins synthesized ...
337-455 6.27e-04

Keratin, high sulfur B2 protein; High sulfur proteins are cysteine-rich proteins synthesized during the differentiation of hair matrix cells, and form hair fibres in association with hair keratin intermediate filaments. This family has been divided up into four regions, with the second region containing 8 copies of a short repeat. This family is also known as B2 or KAP1.


Pssm-ID: 366678 [Multi-domain]  Cd Length: 161  Bit Score: 42.86  E-value: 6.27e-04
                           10        20        30        40        50        60        70        80
                   ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801   337 CEEADCLDPGCSNHGVCihGECHCNPGWGGSNCeILKTMCADQCSGHGTYLQESGSCTCDPNwTGPDCSNEICSVDCGSH 416
Cdd:pfam01500    4 CGTSFCGFPTCSTGGTC--GSGCCQPCCCQSSC-CRPSCCQTSCCQPTTFQSSCCRPTCQPC-CQTSCCQPTCCQTSSCQ 79
                           90       100       110
                   ....*....|....*....|....*....|....*....
gi 568954801   417 GVCMGGSCRCEEGWTGPACNQRACHPRCAEHGTCKDGKC 455
Cdd:pfam01500   80 TGCGGIGYGQEGSSGAVSSRTRWCRPDCRVEGTCLPPCC 118
NHL_like_6 cd14962
Uncharacterized NHL-repeat domain in bacterial proteins; The NHL (NCL-1, HT2A and LIN-41) ...
1027-1246 1.01e-03

Uncharacterized NHL-repeat domain in bacterial proteins; The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.


Pssm-ID: 271332 [Multi-domain]  Cd Length: 271  Bit Score: 43.34  E-value: 1.01e-03
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1027 GKAVEATLMSPKGMAIDKNGLIYFVDGT--MIRKVDQNGIISTLLGSNDLtsarpltcdtsmhisQVRlewPTDLAINPM 1104
Cdd:cd14962    49 GNAGPNRFVSPIGVAIDANGNLYVSDAElgKVFVFDRDGKFLRAIGAGAL---------------FKR---PTGIAVDPA 110
                          90       100       110       120       130       140       150       160
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1105 DNSIYVLDnnvvlqiTENRQVRI--AAGRPMHcQVPgveyPVGKHAVQttLESATAIAVSYSGVLYITETDEKKINRI-- 1180
Cdd:cd14962   111 GKRLYVVD-------TLAHKVKVfdLDGRLLF-DIG----KRGSGPGE--FNLPTDLAVDRDGNLYVTDTMNFRVQIFda 176
                         170       180       190       200       210       220       230       240
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1181 --RQVTTDGEISLVAG---IPSECDCKNDAN---CDCYQS---------------GDGYAKDAKLNAPSSLAASPDGTLY 1237
Cdd:cd14962   177 dgKFLRSFGERGDGPGsfaRPKGIAVDSEGNiyvVDAAFDnvqifnpegellltvGGPGSGPGEFYLPSGIAIDKDDRIY 256

                  ....*....
gi 568954801 1238 IADLGNIRI 1246
Cdd:cd14962   257 VVDQFNRRI 265
EGF_2 pfam07974
EGF-like domain; This family contains EGF domains found in a variety of extracellular proteins.
347-369 1.18e-03

EGF-like domain; This family contains EGF domains found in a variety of extracellular proteins.


Pssm-ID: 400365  Cd Length: 26  Bit Score: 38.10  E-value: 1.18e-03
                           10        20
                   ....*....|....*....|....*
gi 568954801   347 CSNHGVCIH--GECHCNPGWGGSNC 369
Cdd:pfam07974    2 CSGRGTCVNqcGKCVCDSGYQGATC 26
EGF_2 pfam07974
EGF-like domain; This family contains EGF domains found in a variety of extracellular proteins.
380-404 1.23e-03

EGF-like domain; This family contains EGF domains found in a variety of extracellular proteins.


Pssm-ID: 400365  Cd Length: 26  Bit Score: 38.10  E-value: 1.23e-03
                           10        20
                   ....*....|....*....|....*
gi 568954801   380 CSGHGTYLQESGSCTCDPNWTGPDC 404
Cdd:pfam07974    2 CSGRGTCVNQCGKCVCDSGYQGATC 26
YD_repeat_2x TIGR01643
YD repeat (two copies); This model describes two tandem copies of a 21-residue extracellular ...
1574-1614 1.80e-03

YD repeat (two copies); This model describes two tandem copies of a 21-residue extracellular repeat found in Gram-negative, Gram-positive, and animal proteins. The repeat is named for a YD dipeptide, the most strongly conserved motif of the repeat. These repeats appear in general to be involved in binding carbohydrate; the chicken teneurin-1 YD-repeat region has been shown to bind heparin.


Pssm-ID: 273728 [Multi-domain]  Cd Length: 42  Bit Score: 37.95  E-value: 1.80e-03
                           10        20        30        40
                   ....*....|....*....|....*....|....*....|..
gi 568954801  1574 YSSTGQ-IASIQRGTTSEKVDYDSQGRIVSRVFADGKTWSYT 1614
Cdd:TIGR01643    1 YDAAGRlTGSTDADGTTTRYTYDAAGRLVEITDADGGSTRYE 42
YvrE COG3386
Sugar lactone lactonase YvrE [Carbohydrate transport and metabolism]; Sugar lactone lactonase ...
922-1065 2.05e-03

Sugar lactone lactonase YvrE [Carbohydrate transport and metabolism]; Sugar lactone lactonase YvrE is part of the Pathway/BioSystem: Non-phosphorylated Entner-Doudoroff pathway


Pssm-ID: 442613 [Multi-domain]  Cd Length: 266  Bit Score: 42.57  E-value: 2.05e-03
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  922 LACGIDGSLYVGDFNYVR------RIFPSGNVTSVLElrnkDFrHSSN-----PAHRYylatdpvtgdLYVSDTNTRRIY 990
Cdd:COG3386    98 GVVDPDGRLYFTDMGEYLptgalyRVDPDGSLRVLAD----GL-TFPNgiafsPDGRT----------LYVADTGAGRIY 162
                          90       100       110       120       130       140       150
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*...
gi 568954801  991 R-PKSLTGAkdLTkNAEVVAgtgeqclpfdEARCGDGGkaveatlmsPKGMAIDKNGLIY--FVDGTMIRKVDQNGII 1065
Cdd:COG3386   163 RfDLDADGT--LG-NRRVFA----------DLPDGPGG---------PDGLAVDADGNLWvaLWGGGGVVRFDPDGEL 218
Vgb COG4257
Streptogramin lyase [Defense mechanisms];
914-991 2.57e-03

Streptogramin lyase [Defense mechanisms];


Pssm-ID: 443399 [Multi-domain]  Cd Length: 270  Bit Score: 42.31  E-value: 2.57e-03
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801  914 NKLLAPVALACGIDGSLYVGDF--NYVRRIFP-SGNVTSvlelrnkdFRHSSNPAHRYYLATDPvTGDLYVSDTNTRRIY 990
Cdd:COG4257   185 TPGAGPRGLAVDPDGNLWVADTgsGRIGRFDPkTGTVTE--------YPLPGGGARPYGVAVDG-DGRVWFAESGANRIV 255

                  .
gi 568954801  991 R 991
Cdd:COG4257   256 R 256
EGF_CA cd00054
Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular ...
341-370 2.86e-03

Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular (mostly animal) proteins. Many of these proteins require calcium for their biological function and calcium-binding sites have been found to be located at the N-terminus of particular EGF-like domains; calcium-binding may be crucial for numerous protein-protein interactions. Six conserved core cysteines form three disulfide bridges as in non calcium-binding EGF domains, whose structures are very similar. EGF_CA can be found in tandem repeat arrangements.


Pssm-ID: 238011  Cd Length: 38  Bit Score: 37.23  E-value: 2.86e-03
                          10        20        30
                  ....*....|....*....|....*....|....*
gi 568954801  341 DCLDPG-CSNHGVCIHGE----CHCNPGWGGSNCE 370
Cdd:cd00054     4 ECASGNpCQNGGTCVNTVgsyrCSCPPGYTGRNCE 38
EGF_CA cd00054
Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular ...
488-517 3.62e-03

Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular (mostly animal) proteins. Many of these proteins require calcium for their biological function and calcium-binding sites have been found to be located at the N-terminus of particular EGF-like domains; calcium-binding may be crucial for numerous protein-protein interactions. Six conserved core cysteines form three disulfide bridges as in non calcium-binding EGF domains, whose structures are very similar. EGF_CA can be found in tandem repeat arrangements.


Pssm-ID: 238011  Cd Length: 38  Bit Score: 37.23  E-value: 3.62e-03
                          10        20        30
                  ....*....|....*....|....*....|
gi 568954801  488 PGLCNSNGRCTLDQNGWHCVCQPGWRGAGC 517
Cdd:cd00054     8 GNPCQNGGTCVNTVGSYRCSCPPGYTGRNC 37
NHL_like_5 cd14963
Uncharacterized NHL-repeat domain in bacterial proteins; The NHL (NCL-1, HT2A and LIN-41) ...
1221-1317 5.18e-03

Uncharacterized NHL-repeat domain in bacterial proteins; The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.


Pssm-ID: 271333 [Multi-domain]  Cd Length: 268  Bit Score: 41.12  E-value: 5.18e-03
                          10        20        30        40        50        60        70        80
                  ....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 568954801 1221 KLNAPSSLAASPDGTLYIADLGNIRIRAVSKN----------KPLLNSM----------NFYeVASPTDQELYIFDINGT 1280
Cdd:cd14963    54 EFKYPYGIAVDSDGNIYVADLYNGRIQVFDPDgkflkyfpekKDRVKLIspaglaiddgKLY-VSDVKKHKVIVFDLEGK 132
                          90       100       110       120
                  ....*....|....*....|....*....|....*....|..
gi 568954801 1281 HQYTVSLVtGDYLYNFSYSN----DNDVT-AVTDSNGNtlRI 1317
Cdd:cd14963   133 LLLEFGKP-GSEPGELSYPNgiavDEDGNiYVADSGNG--RI 171
RHS_repeat pfam05593
RHS Repeat; RHS proteins contain extended repeat regions. These repeats often appear to be ...
1366-1397 7.04e-03

RHS Repeat; RHS proteins contain extended repeat regions. These repeats often appear to be involved in ligand binding. Note that this model may not find all the repeats in a protein and that it covers two RHS repeats. The 3D structure of an RHS-repeat-containing protein (the B and C components of an ABC toxin complex) has been determined. The RHS repeats form an extended strip of beta-sheet that spirals around to form a hollow shell, encapsulating the variable C-terminal domain.


Pssm-ID: 461685 [Multi-domain]  Cd Length: 37  Bit Score: 36.42  E-value: 7.04e-03
                           10        20        30
                   ....*....|....*....|....*....|..
gi 568954801  1366 GLLATKSDETGWTTFFDYDSEGRLTNVTFPTG 1397
Cdd:pfam05593    5 GRLTSVTDPDGRVTTYTYDAAGRLTAVTDPDG 36
EGF_Tenascin pfam18720
Tenascin EGF domain; This entry represents the EGF-like domains found in tenascin proteins.
315-337 8.29e-03

Tenascin EGF domain; This entry represents the EGF-like domains found in tenascin proteins.


Pssm-ID: 376143  Cd Length: 29  Bit Score: 36.12  E-value: 8.29e-03
                           10        20
                   ....*....|....*....|...
gi 568954801   315 CGGRGICIMGSCACNSGYKGENC 337
Cdd:pfam18720    6 CSSRGVCVDGQCICDSEYSGDDC 28
EGF pfam00008
EGF-like domain; There is no clear separation between noise and signal. pfam00053 is very ...
491-514 9.32e-03

EGF-like domain; There is no clear separation between noise and signal. pfam00053 is very similar, but has 8 instead of 6 conserved cysteines. Includes some cytokine receptors. The EGF domain misses the N-terminus regions of the Ca2+ binding EGF domains (this is the main reason of discrepancy between swiss-prot domain start/end and Pfam). The family is hard to model due to many similar but different sub-types of EGF domains. Pfam certainly misses a number of EGF domains.


Pssm-ID: 394967  Cd Length: 31  Bit Score: 35.82  E-value: 9.32e-03
                           10        20
                   ....*....|....*....|....
gi 568954801   491 CNSNGRCTLDQNGWHCVCQPGWRG 514
Cdd:pfam00008    6 CSNGGTCVDTPGGYTCICPEGYTG 29
YD_repeat_2x TIGR01643
YD repeat (two copies); This model describes two tandem copies of a 21-residue extracellular ...
1362-1400 9.40e-03

YD repeat (two copies); This model describes two tandem copies of a 21-residue extracellular repeat found in Gram-negative, Gram-positive, and animal proteins. The repeat is named for a YD dipeptide, the most strongly conserved motif of the repeat. These repeats appear in general to be involved in binding carbohydrate; the chicken teneurin-1 YD-repeat region has been shown to bind heparin.


Pssm-ID: 273728 [Multi-domain]  Cd Length: 42  Bit Score: 36.03  E-value: 9.40e-03
                           10        20        30
                   ....*....|....*....|....*....|....*....
gi 568954801  1362 HGNSGLLATKSDETGWTTFFDYDSEGRLTNVTFPTGVVT 1400
Cdd:TIGR01643    1 YDAAGRLTGSTDADGTTTRYTYDAAGRLVEITDADGGST 39
 
Blast search parameters
Data Source: Precalculated data, version = cdd.v.3.21
Preset Options:Database: CDSEARCH/cdd   Low complexity filter: no  Composition Based Adjustment: yes   E-value threshold: 0.01

References:

  • Wang J et al. (2023), "The conserved domain database in 2023", Nucleic Acids Res.51(D)384-8.
  • Lu S et al. (2020), "The conserved domain database in 2020", Nucleic Acids Res.48(D)265-8.
  • Marchler-Bauer A et al. (2017), "CDD/SPARCLE: functional classification of proteins via subfamily domain architectures.", Nucleic Acids Res.45(D)200-3.
Help | Disclaimer | Write to the Help Desk
NCBI | NLM | NIH