Files
Python-DA-and-DM/美赛2020C题/Problem_C_Data/hw1 part2/新建文件夹/sentence_splitter_input.txt
T
2020-06-14 13:12:12 +08:00

14 lines
10 KiB
Plaintext
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
Prior to joining the Navy, Hopper earned a Ph.D. in mathematics from Yale University and was a professor of mathematics at Vassar College. Hopper attempted to enlist in the Navy during World War II but was rejected because she was 34 years old. She instead joined the Navy Reserves. Hopper began her computing career in 1944 when she worked on the Harvard Mark I team led by Howard H. Aiken. In 1949, she joined the EckertMauchly Computer Corporation and was part of the team that developed the UNIVAC I computer. At EckertMauchly she began developing the compiler. She believed that a programming language based on English was possible. Her compiler converted English terms into machine code understood by computers.By 1952, Hopper had finished her program linker (originally called a compiler), which was written for the A-0 System. During her wartime service, she co-authored three papers based on her work on the Harvard Mark 1. In 1954, EckertMauchly chose Hopper to lead their department for automatic programming, and she led the release of some of the first compiled languages like FLOW-MATIC. In 1959, she participated in the CODASYL consortium, which consulted Hopper to guide them in creating a machine-independent programming language. This led to the COBOL language, which was inspired by her idea of a language being based on English words.In 1966, she retired from the Naval Reserve, but in 1967 the Navy recalled her to active duty.
A carbohydrate structural variant of MM glycoprotein (glycophorin A). A variant of the MM glycoprotein (glycophorin A) was isolated from erythrocyte membranes of two individual donors, a mother (L.G.) and daughter (V.W.). This glycoprotein was found to be a carbohydrate variant in which, for both donors, certain O-glycosidically linked saccharides retained the core structure consisting of NeuAc(alpha 2,3)Gal(beta 1,3)GalNAc that is common to all O-linked saccharides of the MN glycoproteins, and, in addition, contained substituents, of varying chain lengths, on the primary carbinol of GalNAc. These saccharides were released from the polypeptide by beta-elimination in the presence of sodium borohydride, and aspects of their structure were investigated by glycosidase digestion and periodate oxidation. Thus, the smallest variant structure was deduced to be NeuAc(alpha 2,3)Gal(beta 1,3)[GlcNAc(beta 1,6)]H2GalNAc. The 6-O-linked GlcNAc appears to serve as the focus of further chain elongation reactions, involving alternate additions of Gal and GlcNAc residues and leading to the formation of several homologous structures. Two such structures, NeuAc(alpha 2,3)Gal(beta 1,3)[GlcNAc(beta 1,?) Gal(beta 1,3/4)GlcNAc(beta 1,6)]H2GalNAc and NeuAc(alpha 2,3) Gal(beta 1,3)[Gal(beta 1,3/4)GlcNAc(beta 1,6)]H2GalNAc were the predominant species present. A larger saccharide was also isolated and its partial sequence was determined to be Gal(beta 1,3/4)GlcNAc(beta 1,?)[Gal(beta 1,3/4)Glc-NAc(beta 1,?)] Gal(beta 1,3/4)GlcNAc(beta 1,6)[NeuAc(alpha 2,3)Gal-(beta 1,3)]H2GalNAc. Because the peptide portion of these glycoproteins contains two methionine residues, it was possible to isolate two CNBr glycopeptides from separate regions of the molecule, and to assess the distribution of these variant structures in the polypeptide.
Characterization and Antioxidant Activity Determination of Neutral and Acidic Polysaccharides from Panax Ginseng C. A. Meyer. Panax ginseng (P. ginseng) is the most widely consumed herbal plant in Asia and is well-known for its various pharmacological properties. Many studies have been devoted to this natural product. However, polysaccharide's components of ginseng and their biological effects have not been widely studied. In this study, white ginseng neutral polysaccharide (WGNP) and white ginseng acidic polysaccharide (WGAP) fractions were purified from P. ginseng roots. The chemical properties of WGNP and WGAP were investigated using various chromatography and spectroscopy techniques, including high-performance gel permeation chromatography, Fourier-transform infrared spectroscopy, and high-performance liquid chromatography with an ultra-violet detector. The antioxidant, anti-radical, and hydrogen peroxide scavenging activities were evaluated in vitro and in vivo using Caenorhabditis elegans as the model organism. Our in vitro data by ABTS (2,2'-azino-bis-(3-ethylbenzothiazoline-6-sulfonic acid), reducing power, ferrous ion chelating, and hydroxyl radical scavenging activity suggested that the WGAP with significantly higher uronic acid content and higher molecular weight exhibits a much stronger antioxidant effect as compared to that of WGNP. Similar antioxidant activity of WGAP was also confirmed in vivo by evaluating internal reactive oxygen species (ROS) concentration and lipid peroxidation. In conclusion, WGAP may be used as a natural antioxidant with potent scavenging and metal chelation properties.
Protein Thermodynamic Destabilization in the Assessment of Pathogenicity of a Variant of Uncertain Significance in Cardiac Myosin Binding Protein C. In the era of next generation sequencing (NGS), genetic testing for inherited disorders identifies an ever-increasing number of variants whose pathogenicity remains unclear. These variants of uncertain significance (VUS) limit the reach of genetic testing in clinical practice. The VUS for hypertrophic cardiomyopathy (HCM), the most common familial heart disease, constitute over 60% of entries for missense variants shown in ClinVar database. We have studied a novel VUS (c.1809T>G-p.I603M) in the most frequently mutated gene in HCM, MYBPC3, which codes for cardiac myosin-binding protein C (cMyBPC). Our determinations of pathogenicity integrate bioinformatics evaluation and functional studies of RNA splicing and protein thermodynamic stability. In silico prediction and mRNA analysis indicated no alteration of RNA splicing induced by the variant. At the protein level, the p.I603M mutation maps to the C4 domain of cMyBPC. Although the mutation does not perturb much the overall structure of the C4 domain, the stability of C4 I603M is severely compromised as detected by circular dichroism and differential scanning calorimetry experiments. Taking into account the highly destabilizing effect of the mutation in the structure of C4, we propose reclassification of variant p.I603M as likely pathogenic. Looking into the future, the workflow described here can be used to refine the assignment of pathogenicity of variants of uncertain significance in MYBPC3.
The idea of neural language models as introduced by Bengio et al. [5] is to jointly learn an em- bedding of words into an n-dimensional vector space and to use these vectors to predict how likely a word is given its context. Collobert and Weston [6] introduced a new neural network model to compute such an embedding. When these networks are optimized via gradient ascent the derivatives modify the word embedding matrix L ∈ Rn×|V |, where |V | is the size of the vocabulary. The word vectors inside the embedding matrix capture distributional syntactic and semantic information via the words co-occurrence statistics. For further details and evaluations of these embeddings, see [5, 6, 7, 8]. Once this matrix is learned on an unlabeled corpus, we can use it for subsequent tasks by using each words vector (a column in L) to represent that word. In the remainder of this paper, we represent a sentence (or any n-gram) as an ordered list of these vectors (x1 , . . . , xm ). This word representation is better suited for autoencoders than the binary number representations used in previous related autoencoder models such as the recursive autoassociative memory (RAAM) model of Pollack [9, 10] or recurrent neural networks [11] since the activations are inherently continuous.
The sequences mediating receptor mRNA down-regulation are represented within the AR cDNA and not within the CMV promoter. Androgenic down-regulation of AR cDNA expression was time- and dose-dependent, resembling native AR mRNA down-regulation. In addition, androgenic regulation of the receptor cDNA was not dependent on protein synthesis suggesting that AR and/or another pre-existing protein(s) is involved in this process. In COS 1 cells co-transfected with androgen and glucocorticoid receptor cDNAs, dexamethasone mimicked the action of androgen in down-regulating AR mRNA. This response depended on glucocorticoid receptors. Androgen had little effect on steady-state levels of AR protein consistent with reports that androgen down-regulates AR mRNA but increases AR protein half-life (Kemppainen et al. (1992) J. Biol. Chem. 267, 968-974; Zhou et al. (1995) Mol. Endocrinol. 9, 208-218). However, glucocorticoids decreased AR protein levels in cells that co-expressed androgen and glucocorticoid receptors. These results indicate that sequences represented in the AR cDNA mediate AR mRNA down-regulation by both androgens and glucocorticoids. Inhibition of AR mRNA and protein by glucocorticoids suggests that these steroids may modulate androgen action in tissues, such as mammary gland and prostate, which express both androgen and glucocorticoid receptors.
Immunohistochemical application of antibodies against heparan sulfate proteoglycan core protein and heparitinase-digested heparan sulfate stubs showed the presence of heparan sulfate proteoglycan in all basement membranes of the rat kidney. However, a monoclonal antibody (JM-403) against native heparan sulfate (van den Born, J., van den Heuvel, L. P. W. J., Bakker, M. A. H., Veerkamp, J. H., Assmann, K. J. M., and Berden, J. H. M. (1992) Kidney Int. 41, 115-123) largely failed to stain tubular basement membranes, suggesting the presence of heparan sulfate chains lacking the specific JM-403 epitope. Heparan sulfate preparations from various sources differed markedly with regard to JM-403 binding, as demonstrated by liquid phase inhibition in enzyme-linked immunosorbent assay, the interaction decreasing with increasing sulfate contents of the polysaccharide. Mapping of the JM-403 epitope indicated that it was dominated by one or more N-unsubstituted glucosamine unit(s), since treatments that destroyed or altered the structure of such units in heparan sulfate preparations (cleavage at N-unsubstituted glucosamine units with HNO2 at pH 3.9 and N-acetylation with acetic anhydride, respectively), abolished antibody binding.