Identifying novel transcripts and novel genes in the human genome by using novel SAGE tags

Proc Natl Acad Sci U S A. 2002 Sep 17;99(19):12257-62. doi: 10.1073/pnas.192436499. Epub 2002 Sep 4.

Abstract

The number of genes in the human genome is still a controversial issue. Whereas most of the genes in the human genome are said to have been physically or computationally identified, many short cDNA sequences identified as tags by use of serial analysis of gene expression (SAGE) do not match these genes. By performing experimental verification of more than 1,000 SAGE tags and analyzing 4,285,923 SAGE tags of human origin in the current SAGE database, we examined the nature of the unmatched SAGE tags. Our study shows that most of the unmatched SAGE tags are truly novel SAGE tags that originated from novel transcripts not yet identified in the human genome, including alternatively spliced transcripts from known genes and potential novel genes. Our study indicates that by using novel SAGE tags as probes, we should be able to identify efficiently many novel transcripts/novel genes in the human genome that are difficult to identify by conventional methods.

Publication types

  • Research Support, Non-U.S. Gov't
  • Research Support, U.S. Gov't, P.H.S.

MeSH terms

  • Base Pair Mismatch
  • Base Sequence
  • DNA, Complementary / genetics
  • Databases, Nucleic Acid
  • Expressed Sequence Tags
  • Gene Expression Profiling
  • Genome, Human*
  • Humans
  • Molecular Sequence Data
  • Sequence Tagged Sites*
  • Transcription, Genetic

Substances

  • DNA, Complementary

Associated data

  • GENBANK/BM285378
  • GENBANK/BM285379
  • GENBANK/BM285380
  • GENBANK/BM285381
  • GENBANK/BM285382
  • GENBANK/BM285383
  • GENBANK/BM285384
  • GENBANK/BM285385
  • GENBANK/BM285386
  • GENBANK/BM285387
  • GENBANK/BM285388
  • GENBANK/BM285389
  • GENBANK/BM285390
  • GENBANK/BM285391
  • GENBANK/BM285392
  • GENBANK/BM285393
  • GENBANK/BM285394
  • GENBANK/BQ635328
  • GENBANK/BQ635329
  • GENBANK/BQ635330
  • GENBANK/BQ635331
  • GENBANK/BQ635332