Pseudomonas population genomics: Difference between revisions

From QiuLab
Jump to navigation Jump to search
imported>Rayrah
imported>Rayrah
Line 17: Line 17:
## Protein ortholog data: [http://pseudomonas.com/alignPolymorphicGeneSequencesStep1.do?feature_id_parent=105676&start=1583956&stop=1584798&replicon_id_reference=136&alphabet=protein&limit_to_species=false fleN], [http://pseudomonas.com/alignPolymorphicGeneSequencesStep1.do?feature_id_parent=104960&start=1187587&stop=1189059&replicon_id_reference=136&alphabet=protein&limit_to_species=false fleQ], [http://pseudomonas.com/alignPolymorphicGeneSequencesStep1.do?feature_id_parent=105674&start=1582528&stop=1583817&replicon_id_reference=136&alphabet=protein&limit_to_species=false flhF]     
## Protein ortholog data: [http://pseudomonas.com/alignPolymorphicGeneSequencesStep1.do?feature_id_parent=105676&start=1583956&stop=1584798&replicon_id_reference=136&alphabet=protein&limit_to_species=false fleN], [http://pseudomonas.com/alignPolymorphicGeneSequencesStep1.do?feature_id_parent=104960&start=1187587&stop=1189059&replicon_id_reference=136&alphabet=protein&limit_to_species=false fleQ], [http://pseudomonas.com/alignPolymorphicGeneSequencesStep1.do?feature_id_parent=105674&start=1582528&stop=1583817&replicon_id_reference=136&alphabet=protein&limit_to_species=false flhF]     
## Protocol:
## Protocol:
### Fasta headers are too long for tree, run: [[File:renamer.pl]] editing/rule (script name?); Alignment (tool and command); Phylogeny; Tree display (R commands)
### Fasta headers are too long for tree, run: [https://www.dropbox.com/s/x2p4joeqg7omfub/rename.pl rename.pl] to shorten names <code> usage: >rename.pl FASTA file > OUTPUTfilename.fas</code> editing/rule (script name?); Alignment (tool and command); Phylogeny; Tree display (R commands)
{| class="wikitable"
{| class="wikitable"
! Gene !! Alignment !! Tree !! Notes  
! Gene !! Alignment !! Tree !! Notes  

Revision as of 20:53, 18 June 2013

Projects

  1. Build a local genome database
    1. Database schema:
      1. "genome": genome_id, strain_name, ncbi_taxid
      2. "orf": genome_id, locus_tag, start, stop, strand, genome_name, product_name
      3. "orth_orf": orth_orf_id, locus_name, genome_id, orth_class
    2. Parsing scripts
      1. Rayees Parsing code, requires that you remove columns 9-27 using bash command: cut -c 1-8 (I will write a bash script that does this and runs the program) https://www.dropbox.com/s/lpxxbkxeyw7frrn/parser.pl
    3. Database loading scripts
  2. Molecular Evolution of flagellum genes
    1. Download orthologs
    2. Reconstruct phylogenetic tree
    3. Run PAML tests

Update: June 18, 2013

  1. Material and Methods
    1. Protein ortholog data: fleN, fleQ, flhF
    2. Protocol:
      1. Fasta headers are too long for tree, run: rename.pl to shorten names usage: >rename.pl FASTA file > OUTPUTfilename.fas editing/rule (script name?); Alignment (tool and command); Phylogeny; Tree display (R commands)
Gene Alignment Tree Notes
fleN fleN pep alignment Flen.png Conserved Domain
fleQ fleQ pep alignment Fleq.png Conserved Domain
flhf flhF pep alignment Flhf.png Conserved Domain

Benchmark: June 11, 2013

  1. Finish parsing the genome files to upload the "orf" table (Raymond & Rayees)
    1. Rayees Parsed genome files: https://www.dropbox.com/sh/k0zktvvmv39op9i/1zBercEky8
  2. Parsing the ortholog file to upload the "orth_orf" table (Raymond)
  3. Identify and download fleN, fleQ, and flhF orthologs & align them (Rayees)