Search This Blog

Monday, 25 April 2011

ORganic VIrtual Library (ORVIL) – A combinatorial library construction based on organic constituents and without scaffold hopping


ABSTRACT                                                                                                                                
Rapid construction of virtual combinatorial library is a prerequisite for in silico library enumeration and design. ORganic VIrtual Library (ORVIL) is a perl program to generate a combinatorial library using most frequently observed 200 organic substituents. It is designed to explore the organic chemical space in the given query structure without affecting the entire backbone of the molecule enabling minimum molecular complexities. Its particular features are its simplicity to use, portable SMILES format and high speed of library construction. Benchmarking of Tamoxifen (drug for breast cancer) was performed which revealed a compound having similar architecture of known drug analogue, Toremifene.
Keywords: virtual combinatorial library, organic substituents, molecular complexities, SMILES format

Sunday, 24 April 2011

THE BLESSINGS OF MANAKULA VINAYAGAR


Prasanth Virtual Bioinfo Lab

The blessings of Manakula Vinayagar is always showered on me and I pray my god to bless me forever for success and prosperity in my life and the people around me. Shown image is of Shri Manakula Vinayagar, Pondicherry.

Saturday, 23 April 2011

COMPUTATIONAL STUDIES ON THE INTERACTION OF CORE HISTONE TAIL DOMAINS WITH CpG ISLAND


S. PRASANTH KUMAR 1, RAVI G. KAPOPARA 1,YOGESH T. JASRAI 1 AND RAKESH M. RAWAL 2.

1. Bioinformatics Laboratory, Department of Botany, University School of Sciences,
Ahmedabad- 380 009. 2. Division of Medicinal Chemistry and Pharmacogenomics, Department of Cancer Biology, Cancer & Research Institute (GCRI), Ahmedabad- 380 016.



Download the accepted galley proof here:
http://www.ijpbs.net/vol-3/issue-1/bio/B%20-%2068.pdf
Copyrights Reserved- International Journal of Pharma and Bio Sciences.



ABSTRACT:
It has been elucidated through in vitro studies that core histone tail domains preferentially interact with linker DNA. In the present study, we studied these interaction computationally using molecular docking and isocontour-based electrostatic map approach in order to identify the domains and regions of H3 and H4 tails and DNA contributing for the physical associativeness. We also explored the interaction made by the linker DNA containing methylated CpG dinucleotides (CpG island) with the normal and post-translational modified histone tails. We report that these interactions are electrostatically unfavored if one of the biomolecular partners is methylated thereby negatively charged zones of DNA and histone tails are required to be absent nearby.

KEYWORDS: Core histone tail domain, linker DNA, CpG island, Molecular docking, Isocontour-based
electrostatic potential map.

Here is the snapshot of a study undergone:

COPYRIGHTED IMAGE Prasanth Virtual Bioinfo Lab
Copyrights 2011 Prasanth Virtual bioinfo Lab

Friday, 15 April 2011

UPLOADING SEQUENCES TO THE DATABASES/SEQUENCE SUBMISSIONS

Sequence can be submitted in NCBI GenBank using:
  1. Sequin
  2. BankIt

Sequin

Sequin is a stand-alone software tool developed by the National Center for Biotechnology Information (NCBI) for submitting and updating sequences to the GenBank, EMBL, and DDBJ databases. Sequin has the capacity to handle long sequences and sets of sequences (segmented entries, as well as population, phylogenetic, and mutation studies). It also allows sequence editing and updating, and provides complex annotation capabilities. In addition, Sequin contains a number of built-in validation functions for enhanced quality assurance.

File Formats Accepted

Sequin normally expects to read sequence files in FASTA format. Note that most sequence analysis software packages include FASTA or "raw" as one of the available output formats. Population studies, phylogenetic studies, mutation studies, and environmental samples may be entered in either FASTA format, or in PHYLIP, NEXUS, MACAW, or FASTA+GAP formats if you are submitting an alignment.

Creating a Submission

Sequin is organized into a series of forms for entering submitting authors, entering organism and sequences, entering information such as strain, gene, and protein names, viewing the complete submission, and editing and annotating the submission. The goal is to go quickly from raw sequence data to an assembled record that can be viewed, edited, and submitted to your database of choice.

Submitting Authors Form: The pages in the Submitting Authors form ask you to provide the release date, a working title, names and contact information of submitting authors, and affiliation information.

Submission page: This page asks for a tentative title for a manuscript describing the sequence and will initially mark the manuscript as being unpublished. When the article is published, the database staff will update the sequence record with the new citation. This page also lets you indicate that a record should be held confidential by the database until a specified date, although the preferred policy is to release the record immediately into the public databases. It also contains pages of contact, author and author’s affiliation.

Sequence Format Form: Submission Type: Single Sequence if you have a single contiguous mRNA or genomic DNA sequence.  Segmented Sequence if you have a single collection of non-overlapping, non-contiguous sequences that cover a specified genetic region from a single source. A standard example is a set of genomic DNA sequences that encode exons from a gene along with fragments of their flanking introns. Gapped Sequence if you have a single non-contiguous mRNA or genomic DNA sequence. A gapped sequence contains specified gaps of known or unknown length where the exact nucleotide sequence has not been determined. Sequence Format: FASTA, FASTA+GAP, NEXUS, PHYLIP,etc. Then we have to fill Organism page and Annotation page (this is optional) before final submission. Now, the program will supply an automatic identifier which will be used for deposition in database and for future correspondence.


BankIt

BankIt is a web based tool developed by the National Center for Biotechnology Information (NCBI) for submitting and updating sequences to the GenBank,


Creating a Submission

Contact Information: Name, address, phone number, fax number and email address of the submitter must be entered when registering and submitting for the first time
Release date information: Immediately after it is processed at NCBI or on a date the submitter specifies

Reference information: Sequence authors: names of the researchers who are credited with the sequence Publication information: Unpublished, In-Press, or Published; and applicable citation information (paper's title, authors, journal title, volume, issue, year, pages)

Submission Category and Type: Original sequencing or Third Party Annotation
Single sequence, sequence set (phylogenetic, population, environmental, etc), or batch

Nucleotide sequence(s): Input (cut-and-paste) single or multiple sequences or Upload them as a FASTA file; FASTA files should include organisms in their definition lines
Sequences must be at least 200 nucleotides long (unless they are complete exons, non-coding RNAs (ncRNAs), microsatellites or ancient DNA)

Molecule type: what was sequenced? (genomic DNA, mRNA, genomic RNA, cRNA, etc)
Topology: linear or circular (circular must be complete, such as a complete plasmid)

Organism name, applicable source modifiers, location : Genus and species names (if not previously provided in FASTA file) If name is new or unrecognized, provide best known taxonomic lineage If genus and/or species names are not known, provide most specific name known (for example:Bacillus sp., Uncultured bacterium, Uncultured archaeon) Most complete name for any synthetic vector (for example: Cloning vector pAB234, Transfer vector p789Abc) Source modifiers include: strain, clone, isolate, specimen-voucher, isolation-source, country Location: organelle (mitochondrion, chloroplast, etc); map and/or chromosome

Features of the sequence: Upload files or use input forms to add all applicable features (for example: CDS, gene, rRNA, tRNA, microsatellite, exon, intron)





PATTERN SEARCHING DATABASES

Patterns are regular expressions matching short sequence motifs usually of biological meaning. This pattern serves as discriminators that help to identify a protein’s family e.g. zinc finger binding motif.  Databases which derive patterns from protein superfamily / family are known as Protein Pattern Databases.


Within a single conserved region (motif), the sequence information may be reduced to a consensus expression (a regular expression), often simply referred to as a pattern.

PROSITE

PROSITE is hosted by ExPaSy. PROSITE is an annotated collection of motif descriptors dedicated to the identification of protein families and domains. The motif descriptors used in PROSITE are either patterns or profiles, which are derived from multiple alignments of homologous sequences. This gives to these motif descriptors the notable advantage of identifying distant relationships between sequences that would have passed unnoticed based solely on pairwise sequence alignment.

The core of the PROSITE database is composed of two text files:

• PROSITE.DAT is a computer readable file that contains all the information necessary to programs that make use of PROSITE to scan sequence(s) for the occurrence of patterns or profiles. This file includes, for each of the entry described, statistics on the number of hits obtained while scanning the SWISS-PROT protein database for a pattern or profile. Cross-references to the corresponding SWISS-PROT entries as well as to matched sequences from the PDB 3D-structure database2 are also provided.

• PROSITE.DOC contains textual information that fully documents each pattern or profile.

PROSITE patterns

In some cases the sequence of an unknown protein is too distantly related to any protein of known structure to detect its resemblance by pairwise sequence alignment. However, relationships can be revealed by the occurrence in its sequence of a particular cluster of residue types, which is variously known as a pattern, motif, signature or fingerprint.

These motifs, typically around 10 to 20 amino acids in length, arise because specific residues and regions thought or proved to be important to the biological function of a group of proteins are conserved in both structure and sequence during evolution. These biologically significant regions or residues are generally:
• Enzyme catalytic sites.
• Prostethic group attachment sites (heme, pyridoxal-phosphate, biotin, etc.).
• Amino acids involved in binding a metal ion.
• Cysteines involved in disulphide bonds.
• Regions involved in binding a molecule (ADP/ATP, GDP/GTP, calcium, DNA, etc.) or

As the sequence of biologically meaningful motifs is evolutionarily conserved, a multiple alignment of them can be reduced to a consensus expression called a regular expression or pattern. Each position of such a pattern can be occupied by any residue from a specified set of acceptable residues, and in addition can be repeated a variable number of times within a specified range. At strictly conserved positions only one particular amino acid is accepted, whereas at other positions several amino acids with similar physicochemical properties can be accepted. It is also possible to define which amino acid(s) is(are) incompatible with a given position, and conserved residues can be separated by gaps of variable lengths.



 

BIOINFORMATICS DATA INTEGRATION SYSTEMS/ SEQUENCE RETRIEVAL SYSTEMS

Sequence Retrieval System (SRS)

SRS is a generic bioinformatics data integration software system. Developed initially in the early 1990s as an academic project at the European Molecular Biology Laboratory (EMBL), the system has evolved into a commercial product and is currently sold under license as a stand-alone software product.

SRS uses proprietary parsing techniques largely based on context-free grammars to parse and index flat-file data. A similar system combined with DOM-based processing rules is used to parse and index XML-formatted data. A relational database connector can be used to integrate data stored in relational database systems. SRS provides a unique common interface for accessing heterogeneous data sources and bypass complexities related to the actual format and storage mechanism for the data. SRS can exploit textual references between different databases and pull together data from disparate sources into a unified view.

SRS is designed from the ground up with extensibility and flexibility in mind, in order to cope with the ever-changing list of databases and formats in the bioinformatics world. SRS relies on a mix of database configuration via meta-definitions and hand-crafted parsers to integrate a wide range of database distributions. These meta-definitions are regularly updated and are also available for extension and modification to all users.
A number of similar commercial systems have been developed that replicate the basic functionality of SRS.


Entrez

The Entrez Global Query Cross-Database Search System is a powerful federated search engine, or web portal that allows users to search many discrete health sciences databases at the National Center for Biotechnology Information (NCBI) website. "Entrez" happens to be the second person plural (or formal) form of the French verb "entrer (to enter)", meaning the invitation "Come in!".

Entrez is the text-based search and retrieval system used at NCBI for all of the major databases, including PubMed, Nucleotide and Protein Sequences, Protein Structures, Complete Genomes, Taxonomy, OMIM, and many others. Entrez is at once an indexing and retrieval system, a collection of data from many sources, and an organizing principle for biomedical information.

Entrez Global Query is an integrated search and retrieval system that provides access to all databases simultaneously with a single query string and user interface. Entrez can efficiently retrieve related sequences, structures, and references. The Entrez system can provide views of gene and protein sequences and chromosome maps. Some textbooks are also available online through the Entrez system.

Entrez Nodes Represent Data


An Entrez “node” is a collection of data that is grouped together and indexed together. It is usually referred to as an Entrez database. In the first version of Entrez, there were three nodes: published articles, nucleotide sequences, and protein sequences. Each node represents specific data objects of the same type, e.g., protein sequences, which are each given a unique ID (UID) within that logical Entrez Proteins node. Records in a node may come from a single source (e.g., all published articles are from PubMed) or many sources (e.g., proteins are from translated Gen-Bank sequences, SWISS-PROT, or PIR)


Ensembl

Ensembl is a bioinformatics project to organize biological information around the sequences of large genomes. It is a comprehensive source of stable automatic annotation of individual genomes, and of the synteny and orthology relationships between them. It is also a framework for integration of any biological data that can be mapped onto features derived from the genomic sequence. Ensembl is available as an interactive Web site, a set of flat files, and as a complete, portable open source software system for handling genomes. All data are provided without restriction, and code is freely available. Ensembl’s aims are to continue to “widen” this biological integration to include other model organisms relevant to understanding human biology as they become available; to “deepen” this integration to provide an ever more seamless linkage between equivalent components in different species; and to provide further classification of functional elements in the genome that have been previously elusive.