Cross-reference to related applications
This application is a 35 U.S.C. .sctn.371 National filing of International Application No. PCT/IS2010/050007, filed Jul. 9, 2010, incorporated herein by reference, which claims priority benefit of Iceland patent application No. 8836, filed Jul. 10, 2009.
Background of the invention
Genetic risk is conferred by subtle differences in the sequence of the genome among individuals in a population. The human genome differs between individuals most frequently due to single nucleotide polymorphisms (SNPs), although other variations are also important. SNPs are located on average every 500 base pairs in the human genome. Accordingly, a typical human gene containing 250,000 base pairs may contain approximately 500 different SNPs. Only a minor number of SNPs are located in exons and alter the amino acid sequence of the protein encoded by the gene. Most SNPs may have no known effect on gene function, while others are known to alter transcription, splicing, translation, or stability of the mRNA encoded by the gene. Additional genetic polymorphisms in the human genome are caused by insertions, deletions, translocations, or inversions of either short or long stretches of DNA.
Parent-of-origin effects (POE) are genetic effects that are transmitted from parents to offspring in such a manner that the expression of the phenotype in the offspring depends on whether the transmission originated from the mother or the father. The effect of a sequence variant in the nuclear genome on the phenotype may depend on its parental origin. In one scenario, the effect is due to imprinting, in which an allele is silenced via an epigentic mechanism such as methylation when inherited from one parent and expressed when inherited from the other parent. In general, however, there are three parent-of-origin effects, i.e. those that arise from epigenetic regulation of gene expression (e.g., imprinting), those that arise from effects of intrauterine environment on the development of the fetus and those that arise from genetic variation in the maternally inherited mitochondrial genome.
Diabetes mellitus, often called diabetes, is a metabolic disease wherein carbohydrate utilization is reduced and lipid and protein utilization is enhanced, and is caused by an absolute or relative deficiency of insulin. In the more severe cases, diabetes is characterized by chronic hyperglycemia, glycosuria, water and electrolyte loss, ketoacidosis and coma. Long term complications can include development of both microvascular complications such as neuropathy, retinopathy and nephropathy and macrovascular complications such as myocardial infarction (MI), stroke and peripheral arterial disease (PAD), caused by generalized degenerative changes in large and small blood vessels. The most common form of diabetes is type 2 diabetes (T2D), (also called non-insulin-dependent diabetes) which is characterized by hyperglycemia due to impaired insulin secretion and insulin resistance in target tissues and increased glucose output by the liver. Both genetic and environmental factors contribute to T2D. For example, obesity plays a major role in the development of T2D. Type 1 diabetes is characterized by loss of insulin-producing beta cells in the islets of Langerhans, leading to insulin deficiency, and represents a majority of diabetes cases affecting children.
The prevalence of T2D worldwide is currently 6% but is projected to rise over the next decade (Amos, A. F., McCarty, D. J., Zimmet, P., Diabet Med 14 Suppl 5, S1 (1997)). This increase in prevalence of T2D is attributed to increasing age of the population and rise in obesity. The health implications of T2D are enormous. In 1995, there were 135 million adults with the disease worldwide. It is estimated that close to 300 million will have T2D in the year 2025 (King, H., et al., Diabetes Care, 21(9): 1414-1431 (1998)). The prevalence of T2D in the adult population in Iceland is 2.5% (Vilbergsson, S., et al., Diabet. Med., 14(6): 491-498 (1997)), which means that approximately 5,000 people over the age of 34 in Iceland have T2D.
Many T2D patients suffer serious complications of chronic hyperglycemia including microvascular complications (nephropathy, neuropathy, retinopathy) and accelerated development of cardiovascular disease (including cerebrovascular disease (stroke), myocardial infarction, and peripheral arterial disease) through macrovascular complications.
In fact, the enormous public health burden of diabetes is largely due to the development of vascular complications of the disease. Cardiovascular disease (CVD) is a major complication and the leading cause of premature death among people with diabetes and accounts for over 75% of all deaths among diabetics. Adults with diabetes are two to four times more likely to have heart disease or suffer a stroke than people without diabetes. Approximately 35% of type 1 diabetes patients die from a cardiovascular disease before age 55, illustrating the devastating consequence of the disease through its cardiovascular complications (Krolewski, A. S. et al. N Engl J Med 317:1390-8 (1987)). The overall prevalence of cardiovascular disease is over 55% in adults with diabetes as compared with 2%-4% of the general population (Asley, R. Levy, A. P. Vasc Health Risk Man 1:19-28 (2005)).
Diabetic retionpathy is the cause of blindness in about 5% of blind people worldwide, and almost everyone with diabetes has some degree of retinopathy after 20 years with the disease (Marshall, S. M. Flyvbjerg, A. British Med J 333:475-80 (2006)). The prevalence of retinopathy is highest in young-onset patients, and steadily increase with duration of diabetes (Chiarelli, F., et al. Horm Res 57(suppl 1):113-6 (2002)).
Nephropathy is also common in diabetic patients, which confers increased risk of premature death due to end-stage renal failure and cardiovascular disease. About half of diabetic patients develop microalbuminuria, which is a marker for early nephropathy, at some point, and about one third will progress to proteinuria. Once present, proteinuria will inevitably lead to end stage renal disease; between 20% and 50% of patients who start renal replacement therapy have diabetes (Marshall, S. M. Flyvbjerg, A. British Med J 333:475-80 (2006)). Patients with diabetes have between 30% and 50% lifetime risk of developing chronic peripheral neuropathy, which can lead to severe symptoms such as foot ulcerations and amputation of lower limbs.
Many of the complications of diabetes have a prolonged subclinical asymptomatic phase. Thus, screening for presymptomatic complications, such as retinopathy and microalbuminuria is extremely important for effective disease management. For example, the micro- and macrovascular complications of diabetes are almost unknown in younger children and rare in adolescents and young adults, but can be detected as soon as 2-5 years after diagnosis during childhood and adolescence (Clarke B. F., in Diabetes Mellitus in Children and Adolescents, Kelnar, C. (ed); London, Chapman & Hall, pp 539-51 (1994)).
As genetic polymorphisms conferring risk of common diseases, such as Type 1 and Type 2 diabetes mellitus, are uncovered, genetic testing for such risk factors is becoming important for clinical medicine. Established examples include apolipoprotein E testing to identify genetic carriers of the apoE4 polymorphism in dementia patients for the differential diagnosis of Alzheimer's T2D, and of Factor V Leiden testing for predisposition to deep venous thrombosis. More importantly, in the treatment of cancer, diagnosis of genetic variants in tumor cells is used for the selection of the most appropriate treatment regime for the individual patient. In breast cancer, genetic variation in estrogen receptor expression or heregulin type 2 (Her2) receptor tyrosine kinase expression determine if anti-estrogenic drugs (tamoxifen) or anti-Her2 antibody (Herceptin) will be incorporated into the treatment plan. In chronic myeloid leukemia (CML) diagnosis of the Philadelphia chromosome genetic translocation fusing the genes encoding the Bcr and Abl receptor tyrosine kinases indicates that Gleevec (STI571), a specific inhibitor of the Bcr-Abl kinase should be used for treatment of the cancer. For CML patients with such a genetic alteration, inhibition of the Bcr-Abl kinase leads to rapid elimination of the tumor cells and remission from leukemia.
Until recently, two approaches were mainly used to search for genes associated with T2D. Single nucleotide polymorphisms (SNPs) within candidate genes have been tested for association and two variants conferring a modest risk of T2D were identified by this method; a protective Pro12Ala polymorphism in the peroxisome proliferator activated receptor gamma gene (PPARG2) (Altshuler, D. et al., Nat Genet. 26, 76 (2000)) and a polymorphism in the potassium inwardly-rectifying channel, subfamily J, member 11 gene (KCNJ11) (Gloyn A. L. et al., Diabetes 52, 568 (2003)). Genome-wide linkage scans in families with the common form of T2D have yielded several loci but the responsible genes within these loci have mostly yet to be uncovered. The rare Mendelian forms of T2D, namely maturity-onset diabetes of the young (MODY), have yielded six genes by positional cloning (Gloyn, A. L., Ageing Res Rev 2, 111 (2003)).
Genome-wide linkage scan for T2D in the Icelandic population showed suggestive evidence of linkage to chromosome 10q (Reynisdottir, I. et al., Am J Hum Genet. 73, 323 (2003)). Fine mapping of this locus revealed the transcription factor 7-like 2 gene (TCF7L2; formerly TCF4) as being associated with T2D (P=2.1.times.10(-9)) (Grant, S. F. et al., Nat Genet. 38, 320 (2006)). Compared with non-carriers, heterozygous and homozygous carriers of the at-risk alleles (38% and 7% of the population, respectively) have relative risks of 1.45 and 2.41. This corresponds to a population attributable risk of 21%. Association of the TCF7L2 variant has now been replicated in a large number of independent studies with similar relative risk found in the different populations studied. The TCF7L2 gene product is a high mobility group box-containing transcription factor previously implicated in blood glucose homeostasis. It is thought to act through regulation of proglucagon gene expression in enteroendocrine cells via the Wnt signaling pathway.
Recently, genome wide association studies using a large number (300,000-1,000,000) of SNPs have been applied to T2D (Sladek, R et al. Nature. 2007; 445:828-30; Steinthorsdottir V et al. Nat. Gen. 2007; 39:770-5; Saxena, R et al. Science 2007; 316:1331-6; Zeggini, E et al. Science 2007; 316:1336-41; Scott, L J et al. Science 2007; 316:1341-5; Zeggini, E et al. Nat. Gen. 40:638-45 (2008). In addition to confirming the three previously identified variants (PPARG, KCNJ11 and TCF7L2) these studies have thus far identified 11 additional genetic variants conferring risk of T2D. All the variants have a modest risk with TCF7L2 conferring the highest risk. Most, if not all, genome wide studies published to date treat the paternal and maternal alleles as interchangeable. This is likely due to the fact that unless the parents of a proband have been genotyped, the information required to determine the parental origin of alleles is unavailable.
Despite the advances in unraveling the genetics of T2D, the pathophysiology of the T2D remains elusive. However, with the current genetic information we are in a better position to test the effect of different treatment options in relation to the genetic background. It has already been shown that the TCF7L2 at-risk genotype affects the treatment outcome both from lifestyle changes and medication (Florez J C et al. N Engl J Med 2006; 355:241-50; Pearson E R et al. Diabetes 2007; 2178-82).
While our understanding of the genetic bases of developing T2D has increased, the genetics of the disease are still not fully explained. There is therefore an unmet medical need to define additional genetic risk factors affecting the development of T2D. Such information could then be used for diagnostic applications, including applications for identifying those at particularly high risk of developing T2D, development of risk management methods, and for risk stratification where individuals at high risk would be targeted for stringent treatment of other risk factors such as glycemia, high cholesterol and hypertension.
Summary of the invention
The present invention relates to materials and methods for predicting disease risk, by determining the parental origin of particular alleles at polymorphic sites. Certain markers have been found to be predictive of risk of certain diseases, including type 2 diabetes, breast cancer and basal cell carcinoma. Such markers are useful in various diagnostic applications, as described further herein.
In a general sense, the invention provides methods of determining susceptibility to a medical condition for a human subject. To determine such susceptibility, sequence information about particular polymorphic markers is obtained. Preferably, the information includes parental origin of particular alleles, and susceptibility to the condition determined based on such information.
In a first aspect the invention provides a method of determining a susceptibility to type 2 diabetes in a human individual, the method comprising (i) obtaining nucleic acid sequence data about a human individual identifying at least one allele of at least one polymorphic marker, and (ii) determining a susceptibility to type 2 diabetes from the sequence data, wherein the at least one polymorphic marker is selected from the group consisting of rs2334499, and markers in linkage disequilibrium therewith.
Another aspect provides a method of determining a susceptibility to type 2 diabetes in a human individual, the method comprising (i) analyzing nucleic acid sequence data from a human individual for at least one polymorphic marker selected from the group consisting of rs2334499, and markers in linkage disequilibrium therewith, and (ii) determining a susceptibility to type 2 diabetes from the nucleic acid sequence data.
The method may include a further step of determining the parental origin of the at least one allele of the at least one polymorphic marker, wherein different parental origins of the at least one allele are associated with different susceptibilities to type 2 diabetes in humans, and determining a susceptibility to type 2 diabetes based on the parental origin of said at least one allele.
In certain embodiments, the at least one polymorphic marker is selected from the group consisting of rs2334499, rs1038727, rs7131362, rs748541, rs4752779, rs4752780, rs4752781, rs4417225, rs10769560, rs17245346, rs11607954, rs10839220, and rs11600502.
In one embodiment, determination of a paternal origin of the T allele of rs2334499, or a marker allele in linkage disequilibrium therewith, is indicative of increased susceptibility of type 2 diabetes in the individual. Further, determination of a maternal origin of the T allele of rs2334499, or a marker allele in linkage disequilibrium therewith, is indicative of a decreased susceptibility of type 2 diabetes in the individual.
Some embodiments include a further step comprising determining whether at least one additional at-risk variant of type 2 diabetes is present in the individual. The at least one at-risk variant is in some embodiments selected from the group consisting of allele T of rs7903146, allele C of rs1801282, allele G of rs7756992, allele T of rs10811661, allele C of rs1111875, allele T of rs4402960, allele T of rs5219, allele C of rs9300039, allele A of rs8050136, allele C of rs13266634, allele T of rs7836388, allele A of rs11775310, allele C of rs1515018, allele C of rs1470579, and allele C of rs7754840.
Certain embodiments further include a step of determining at least one biomarker in the human individual.
Another aspect of the invention relates to a method of determining a susceptibility to type 2 diabetes in a human individual, the method comprising (i) obtaining sequence data about a human individual identifying at least one allele of at least one polymorphic marker, wherein different parental origins of the at least one allele are associated with different susceptibilities to type 2 diabetes in humans; (ii) determining the parental origin of said at least one allele; and (iii) determining a susceptibility to type 2 diabetes for the individual based on the parental origin of said at least one allele; wherein the at least one polymorphic marker is selected from the group consisting of rs2237892, rs231362, rs4731702 and rs2334499, and markers in linkage disequilibrium therewith.
In certain embodiments, determination of a maternal origin of the C allele of rs2237892, a maternal origin of the C allele of rs231362, a maternal origin of the C allele of rs4731702, or a paternal origin of the T allele of rs2334499, or a marker allele in linkage disequilibrium therewith, is indicative of increased susceptibility of type 2 diabetes in the individual.
Also provided is a method of determining a susceptibility to breast cancer in a human individual, the method comprising (i) obtaining sequence data about a human individual identifying at least one allele of at least one polymorphic marker, wherein different parental origins of the at least one allele are associated with different susceptibilities to breast cancer in humans; (ii) determining the parental origin of said at least one allele; and (iii) determining a susceptibility to breast cancer for the individual based on the parental origin of said at least one allele; wherein the at least one polymorphic marker is selected from the group consisting of rs3817198, and markers in linkage disequilibrium therewith. In one embodiment, determination of a paternal origin of the C allele of rs3817198, or a marker allele in linkage disequilibrium therewith, is indicative of increased susceptibility to breast cancer in the individual.
The invention also provides a method of determining a susceptibility to basal cell carcinoma in a human individual, the method comprising (i) obtaining sequence data about a human individual identifying at least one allele of at least one polymorphic marker, wherein different parental origins of the at least one allele are associated with different susceptibilities to basal cell carcinoma in humans; (ii) determining the parental origin of said at least one allele; and (iii) determining a susceptibility to basal cell carcinoma for the individual based on the parental origin of said at least one allele; wherein the at least one polymorphic marker is selected from the group consisting of rs157935, and markers in linkage disequilibrium therewith. In one embodiment, determination of a paternal origin of the T allele of rs157935 is indicative of increased susceptibility to basal cell carcinoma in the individual.
Another aspect of the invention relates to a method of identification of a marker for use in assessing susceptibility to type 2 diabetes, the method comprising (i) identifying at least one polymorphic marker in linkage disequilibrium with at least one of the markers rs2237892, rs231362, rs4731702 and rs2334499; (ii) determining the genotype status of a sample of individuals diagnosed with, or having a susceptibility to, type 2 diabetes; and (iii) determining the genotype status of a sample of control individuals; wherein a significant difference in frequency of at least one allele in at least one polymorphism in individuals diagnosed with, or having a susceptibility to, type 2 diabetes, as compared with the frequency of the at least one allele in the control sample is indicative of the at least one polymorphism being useful for assessing susceptibility to type 2 diabetes.
Determination of an increase in frequency of the at least one allele in the at least one polymorphism in individuals diagnosed with, or having a susceptibility to, type 2 diabetes, as compared with the frequency of the at least one allele in the control sample is in certain embodiments, indicative of the at least one polymorphism being useful for assessing increased susceptibility to type 2 diabetes; and a decrease in frequency of the at least one allele in the at least one polymorphism in individuals diagnosed with, or having a susceptibility to, type 2 diabetes, as compared with the frequency of the at least one allele in the control sample is indicative of the at least one polymorphism being useful for assessing decreased susceptibility to, or protection against, type 2 diabetes.
Also provided is a method of predicting prognosis of a human individual experiencing symptoms associated with, or an individual diagnosed with, type 2 diabetes, the method comprising (i) obtaining sequence information about the human individual identifying at least one allele of at least one polymorphic marker selected from the group consisting of rs2237892, rs231362, rs4731702 and rs2334499, and markers in linkage disequilibrium therewith, wherein different alleles of the at least one polymorphic marker are associated with different susceptibilities to type 2 diabetes in humans, and predicting prognosis of type 2 diabetes of the human individual from the sequence data.
Further provided is a method of assessing an individual for probability of response to a therapeutic agent for preventing, treating and/or ameliorating symptoms associated with type 2 diabetes, comprising (i) obtaining sequence information about the human individual identifying at least one allele of at least one polymorphic marker selected from the group consisting of rs2237892, rs231362, rs4731702 and rs2334499, and markers in linkage disequilibrium therewith, wherein the at least one allele is associated with a probability of a positive response to the therapeutic agent in humans, and determining the probability of a positive response to the therapeutic agent from the sequence data. In certain embodiments, the therapeutic agent is selected from the group consisting of the agents set forth in Agent Table 1 and Agent Table 2.
The invention also provides kits. In one such aspect, a kit for assessing susceptibility to type 2 diabetes in a human individual is provided, the kit comprising (i) reagents for selectively detecting at least one allele of at least one polymorphic marker in the genome of the individual, wherein the polymorphic marker is selected from the group consisting of rs2237892, rs231362, rs4731702 and rs2334499, and markers in linkage disequilibrium therewith, and (ii) a collection of data comprising correlation data between the polymorphic markers assessed by the kit and susceptibility to type 2 diabetes.
Yet another aspect of the invention relates to the use of an oligonucleotide probe in the manufacture of a diagnostic reagent for diagnosing and/or assessing susceptibility to type 2 diabetes in a human individual, wherein the probe is capable of hybridizing to a segment of a nucleic acid whose sequence is given by any one of SEQ ID NO:1-7, wherein the segment is 15-500 nucleotides in length. In a preferred embodiment, the segment of the nucleic acid to which the probe hybridizes comprises a polymorphic site.
Computer-implemented aspects are also provided. One such aspect relates to a computer-readable medium having computer executable instructions for determining susceptibility to type 2 diabetes in a human individual, the computer readable medium comprising (i) data indicative of at least one polymorphic marker; and (ii) a routine stored on the computer readable medium and adapted to be executed by a processor to determine risk of developing type 2 diabetes in an individual for the at least one polymorphic marker; wherein the at least one polymorphic marker is selected from the group consisting of rs2237892, rs231362, rs4731702 and rs2334499, and markers in linkage disequilibrium therewith.
Another such aspect relates to an apparatus for determining a genetic indicator for type 2 diabetes in a human individual, comprising (i) a processor; and (ii) a computer readable memory having computer executable instructions adapted to be executed on the processor to analyze marker and/or haplotype information for at least one human individual with respect to at least one polymorphic marker selected from the group consisting of rs2237892, rs231362, rs4731702 and rs2334499, and markers in linkage disequilibrium therewith, and generate an output based on the marker or haplotype information, wherein the output comprises a risk measure of the at least one marker or haplotype as a genetic indicator of type 2 diabetes for the human individual.
In one embodiment, the computer readable memory further comprises data indicative of the risk of developing diabetes mellitus associated with at least one allele of at least one polymorphic marker or at least one haplotype, and wherein a risk measure for the human individual is based on a comparison of the at least one marker and/or haplotype status for the human individual to the risk of diabetes mellitus associated with the at least one allele of the at least one polymorphic marker or the at least one haplotype.
In another embodiment, the computer readable memory further comprises data indicative_of the frequency of at least one allele of at least one polymorphic marker or at least one haplotype in a plurality of individuals diagnosed with diabetes mellitus, and data indicative of the frequency of at the least one allele of at least one polymorphic marker or at least one haplotype in a plurality of reference individuals, and wherein risk of developing diabetes mellitus is based on a comparison of the frequency of the at least one allele or haplotype in individuals diagnosed with diabetes mellitus and reference individuals.
It should be understood that all combinations of features described herein are contemplated, even if the combination of feature is not specifically found in the same sentence or paragraph herein. This includes in particular the use of all markers disclosed herein, alone or in combination, for analysis individually or in haplotypes, in all aspects of the invention as described herein.
Brief description of the drawings
The foregoing and other objects, features and advantages of the invention will be apparent from the following more particular description of preferred embodiments of the invention.
FIG. 1 provides a diagram illustrating a computer-implemented system utilizing risk variants as described herein.
FIG. 2 shows a diagram of the chromosome 11p15 locus, illustrating the position of the markers rs2334499, rs3817198, rs231362 and rs2237892 relative to genes in the region.
FIG. 3 shows a diagram of the chromosome 7q32 region.
FIG. 4 shows the relative position of the CTCF motif on chromosome 11p15 with respect to rs2334499.
FIG. 5 shows the position on chromosome 11p15 containing a structural polymorphism, and its relationship to rs2334499, the CTCF motif, duplications in the region, and BamHI and HindIII restriction maps (upper half); and a restriction fragment illustrating the polymorphism in 24 individuals (lower half).
Detailed description
Definitions
Unless otherwise indicated, nucleic acid sequences are written left to right in a 5' to 3' orientation. Numeric ranges recited within the specification are inclusive of the numbers defining the range and include each integer or any non-integer fraction within the defined range. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by the ordinary person skilled in the art to which the invention pertains.
The following terms shall, in the present context, have the meaning as indicated:
A "polymorphic marker", sometime referred to as a "marker", as described herein, refers to a genomic polymorphic site. Each polymorphic marker has at least two sequence variations characteristic of particular alleles at the polymorphic site. Thus, genetic association to a polymorphic marker implies that there is association to at least one specific allele of that particular polymorphic marker. The marker can comprise any allele of any variant type found in the genome, including SNPs, mini- or microsatellites, translocations and copy number variations (insertions, deletions, duplications). Polymorphic markers can be of any measurable frequency in the population. For mapping of disease genes, polymorphic markers with population frequency higher than 5-10% are in general most useful. However, polymorphic markers may also have lower population frequencies, such as 1-5% frequency, or even lower frequency, in particular copy number variations (CNV5). The term shall, in the present context, be taken to include polymorphic markers with any population frequency.
An "allele" refers to the nucleotide sequence of a given locus (position) on a chromosome. A polymorphic marker allele thus refers to the composition (i.e., sequence) of the marker on a chromosome. Genomic DNA from an individual contains two alleles (e.g., allele-specific sequences) for any given polymorphic marker, representative of each copy of the marker on each chromosome. Sequence codes for nucleotides used herein are: A=1, C=2, G=3, T=4. For microsatellite alleles, the CEPH sample (Centre d'Etudes du Polymorphisme Humain, genomics repository, CEPH sample 1347-02) is used as a reference, the shorter allele of each microsatellite in this sample is set as 0 and all other alleles in other samples are numbered in relation to this reference. Thus, e.g., allele 1 is 1 bp longer than the shorter allele in the CEPH sample, allele 2 is 2 bp longer than the shorter allele in the CEPH sample, allele 3 is 3 bp longer than the lower allele in the CEPH sample, etc., and allele -1 is 1 bp shorter than the shorter allele in the CEPH sample, allele -2 is 2 bp shorter than the shorter allele in the CEPH sample, etc.
Sequence conucleotide ambiguity as described herein and in the accompanying sequence listing is as proposed by IUPAC-IUB. These codes are compatible with the codes used by the EMBL, GenBank, and PIR databases.
TABLE-US-00001 IUB code Meaning A Adenosine C Cytidine G Guanine T Thymidine R G or A Y T or C K G or T M A or C S G or C W A or T B C, G or T D A, G or T H A, C or T V A, C or G N A, C, G or T (Any base)
A nucleotide position at which more than one sequence is possible in a population (either a natural population or a synthetic population, e.g., a library of synthetic molecules) is referred to herein as a "polymorphic site".
A "Single Nucleotide Polymorphism" or "SNP" is a DNA sequence variation occurring when a single nucleotide at a specific location in the genome differs between members of a species or between paired chromosomes in an individual. Most SNP polymorphisms have two alleles. Each individual is in this instance either homozygous for one allele of the polymorphism (i.e. both chromosomal copies of the individual have the same nucleotide at the SNP location), or the individual is heterozygous (i.e. the two sister chromosomes of the individual contain different nucleotides). The SNP nomenclature as reported herein refers to the official Reference SNP (rs) ID identification tag as assigned to each unique SNP by the National Center for Biotechnological Information (NCBI).
A "variant", as described herein, refers to a segment of DNA that differs from the reference DNA. A "marker" or a "polymorphic marker", as defined herein, is a variant. Alleles that differ from the reference are referred to as "variant" alleles.
A "microsatellite" is a polymorphic marker that has multiple small repeats of bases that are 2-8 nucleotides in length (such as CA repeats) at a particular site, in which the number of repeat lengths varies in the general population. An "indel" is a common form of polymorphism comprising a small insertion or deletion that is typically only a few nucleotides long.
A "haplotype," as described herein, refers to a segment of genomic DNA that is characterized by a specific combination of alleles arranged along the segment. For diploid organisms such as humans, a haplotype comprises one member of the pair of alleles for each polymorphic marker or locus along the segment. In a certain embodiment, the haplotype can comprise two or more alleles, three or more alleles, four or more alleles, or five or more alleles. Haplotypes are described herein in the context of the marker name and the allele of the marker in that haplotype, e.g., "T rs2334499" refers to the 4 allele of marker rs2334499 being in the haplotype, and is equivalent to "rs2334499 allele 4". Furthermore, allelic codes in haplotypes are as for individual markers, i.e. 1=A, 2=C, 3=G and 4=T.
The term "susceptibility", as described herein, refers to the proneness of an individual towards the development of a certain state (e.g., a certain trait, phenotype or disease), or towards being less able to resist a particular state than the average individual. The term encompasses both increased susceptibility and decreased susceptibility. Thus, particular alleles at polymorphic markers and/or haplotypes of the invention as described herein may be characteristic of increased susceptibility (i.e., increased risk) of type 2 diabetes, as characterized by a relative risk (RR) or odds ratio (OR) of greater than one for the particular allele or haplotype. Alternatively, the markers and/or haplotypes of the invention are characteristic of decreased susceptibility (i.e., decreased risk) of type 2 diabetes, as characterized by a relative risk of less than one.
The term "and/or" shall in the present context be understood to indicate that either or both of the items connected by it are involved. In other words, the term herein shall be taken to mean "one or the other or both".
The term "look-up table", as described herein, is a table that correlates one form of data to another form, or one or more forms of data to a predicted outcome to which the data is relevant, such as phenotype or trait. For example, a look-up table can comprise a correlation between allelic data for at least one polymorphic marker and a particular trait or phenotype, such as a particular disease diagnosis, that an individual who comprises the particular allelic data is likely to display, or is more likely to display than individuals who do not comprise the particular allelic data. Look-up tables can be multidimensional, i.e. they can contain information about multiple alleles for single markers simultaneously, or they can contain information about multiple markers, and they may also comprise other factors, such as particulars about diseases, diagnoses, racial information, biomarkers, biochemical measurements, therapeutic methods or drugs, etc.
A "computer-readable medium", is an information storage medium that can be accessed by a computer using a commercially available or custom-made interface. Exemplary computer-readable media include memory (e.g., RAM, ROM, flash memory, etc.), optical storage media (e.g., CD-ROM), magnetic storage media (e.g., computer hard drives, floppy disks, etc.), punch cards, or other commercially available media. Information may be transferred between a system of interest and a medium, between computers, or between computers and the computer-readable medium for storage or access of stored information. Such transmission can be electrical, or by other available methods, such as IR links, wireless connections, etc.
A "nucleic acid sample" as described herein, refers to a sample obtained from an individual that contains nucleic acid (DNA or RNA). In certain embodiments, i.e. the detection of specific polymorphic markers and/or haplotypes, the nucleic acid sample comprises genomic DNA. Such a nucleic acid sample can be obtained from any source that contains genomic DNA, including a blood sample, sample of amniotic fluid, sample of cerebrospinal fluid, or tissue sample from skin, muscle, buccal or conjunctival mucosa, placenta, gastrointestinal tract or other organs.
The term "therapeutic agent for type 2 diabetes" refers to an agent that can be used to ameliorate or prevent symptoms associated with type 2 diabetes.
The term "type 2 diabetes-associated nucleic acid", as described herein, refers to a nucleic acid that has been found to be associated to type 2 diabetes. This includes, but is not limited to, the markers and haplotypes described herein and markers and haplotypes in strong linkage disequilibrium (LD) therewith. In one embodiment, a type 2 diabetes-associated nucleic acid refers to an LD-block found to be associated with Type 2 diabetes through at least one polymorphic marker located within the LD block.
The term "antisense agent" or "antisense oligonucleotide" refers, as described herein, to molecules, or compositions comprising molecules, which include a sequence of purine an pyrimidine heterocyclic bases, supported by a backbone, which are effective to hydrogen bond to a corresponding contiguous bases in a target nucleic acid sequence. The backbone is composed of subunit backbone moieties supporting the purine and pyrimidine heterocyclic bases at positions which allow such hydrogen bonding. These backbone moieties are cyclic moieties of 5 to 7 atoms in size, linked together by phosphorous-containing linkage units of one to three atoms in length. In certain preferred embodiments, the antisense agent comprises an oligonucleotide molecule.
The term "LD Block C11", as described herein, refers to the genomic segment on chromosome 11 between position 1,625,434 and 1,672,208 (inclusive) in the human genome assembly Build 36. The segment has sequence as set forth in SEQ ID NO:7 herein.
Identification of Susceptibility Variants for Type 2 Diabetes
The present inventors have discovered that certain genetic variants confer increased risk of type 2 diabetes. A search for variants associated with type 2 diabetes has revealed that markers in several genomic locations are associated with risk of type 2 diabetes. The inventors have also discovered that certain variants confer risk of breast cancer and basal cell carcinoma. In all cases, the effect of the associated markers is through a mechanism that depends on the parental origin of the associated allele. In other words, the effect is dependent on the parental origin of the associated allele.
Chromosome 11p15 Locus
An association with type 2 diabetes was observed in two distinct regions of chromosome 11p15. Marker rs231362 has previously been reported to be associated with type 2 diabetes. The present inventors have surprisingly found that maternal transmission of the C allele of this marker is associated with increased risk of type 2 diabetes. The present inventors have also surprisingly discovered another variant, rs2334499, in the chromosome 11p15 region that is associated with risk of type 2 diabetes. The association of this marker is striking in that a paternal transmission of the T allele is associated with increased risk of type 2 diabetes, while a maternal transmission of the same allele is associated with a decreased risk of type 2 diabetes. The observed overall risk for the marker, ignoring these parent-of-origin effects, is thus an average of these underlying effects.
The description continues in the full USPTO document.