<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://wiki.genometracker.org/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Wikiuser</id>
	<title>QiuLab - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://wiki.genometracker.org/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Wikiuser"/>
	<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/w/Special:Contributions/Wikiuser"/>
	<updated>2026-07-28T21:34:35Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.39.2</generator>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6748</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6748"/>
		<updated>2026-07-27T07:23:05Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 9 (7/30) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;[https://edu.gitee.com/huntercollege/courses/5124/noticeboard Course link on gitee]&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1, review&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* Linux Tutorial I. Basics. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
* Linux Tutorial II. Advanced&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 4. Linux Intermezzos 1-4&lt;br /&gt;
&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* Linux Tutorial III: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
* Probability by Monte Carlo simulations&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 5. BbWrapper utilities&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2. Linux commands &amp;amp; BpWraper genomics utilities&lt;br /&gt;
* The Birthday Problem&lt;br /&gt;
** Probability by Monte Carlo simulation&lt;br /&gt;
** Simulation algorithms&lt;br /&gt;
** Implementation in R or Python&lt;br /&gt;
** Visualization&lt;br /&gt;
&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
* Random DNA sequences&lt;br /&gt;
* DNA replication&lt;br /&gt;
* DNA transcription&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 6. Implement algorithms &amp;amp; make visualization: B-Day problem; DNA replication &amp;amp; transcription&lt;br /&gt;
&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* 6-frame translation&lt;br /&gt;
** Algorithm&lt;br /&gt;
** Visualization&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 7. Implement &amp;amp; Visualize 6-frame translation&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3: Central Dogma&lt;br /&gt;
&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6747</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6747"/>
		<updated>2026-07-27T07:22:25Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 8 (7/28) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;[https://edu.gitee.com/huntercollege/courses/5124/noticeboard Course link on gitee]&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1, review&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* Linux Tutorial I. Basics. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
* Linux Tutorial II. Advanced&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 4. Linux Intermezzos 1-4&lt;br /&gt;
&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* Linux Tutorial III: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
* Probability by Monte Carlo simulations&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 5. BbWrapper utilities&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2. Linux commands &amp;amp; BpWraper genomics utilities&lt;br /&gt;
* The Birthday Problem&lt;br /&gt;
** Probability by Monte Carlo simulation&lt;br /&gt;
** Simulation algorithms&lt;br /&gt;
** Implementation in R or Python&lt;br /&gt;
** Visualization&lt;br /&gt;
&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
* Random DNA sequences&lt;br /&gt;
* DNA replication&lt;br /&gt;
* DNA transcription&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 6. Implement algorithms &amp;amp; make visualization: B-Day problem; DNA replication &amp;amp; transcription&lt;br /&gt;
&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* 6-frame translation&lt;br /&gt;
** Algorithm&lt;br /&gt;
** Visualization&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3: Central Dogma&lt;br /&gt;
&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6746</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6746"/>
		<updated>2026-07-27T07:22:04Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 8 (7/28) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;[https://edu.gitee.com/huntercollege/courses/5124/noticeboard Course link on gitee]&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1, review&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* Linux Tutorial I. Basics. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
* Linux Tutorial II. Advanced&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 4. Linux Intermezzos 1-4&lt;br /&gt;
&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* Linux Tutorial III: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
* Probability by Monte Carlo simulations&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 5. BbWrapper utilities&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2. Linux commands &amp;amp; BpWraper genomics utilities&lt;br /&gt;
* The Birthday Problem&lt;br /&gt;
** Probability by Monte Carlo simulation&lt;br /&gt;
** Simulation algorithms&lt;br /&gt;
** Implementation in R or Python&lt;br /&gt;
** Visualization&lt;br /&gt;
&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
* Random DNA sequences&lt;br /&gt;
* DNA replication&lt;br /&gt;
* DNA transcription&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 7. Implement algorithms &amp;amp; make visualization: B-Day problem; DNA replication &amp;amp; transcription&lt;br /&gt;
&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* 6-frame translation&lt;br /&gt;
** Algorithm&lt;br /&gt;
** Visualization&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3: Central Dogma&lt;br /&gt;
&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6745</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6745"/>
		<updated>2026-07-27T07:20:32Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 6 (7/23) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;[https://edu.gitee.com/huntercollege/courses/5124/noticeboard Course link on gitee]&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1, review&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* Linux Tutorial I. Basics. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
* Linux Tutorial II. Advanced&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 4. Linux Intermezzos 1-4&lt;br /&gt;
&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* Linux Tutorial III: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
* Probability by Monte Carlo simulations&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 5. BbWrapper utilities&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2. Linux commands &amp;amp; BpWraper genomics utilities&lt;br /&gt;
* The Birthday Problem&lt;br /&gt;
** Probability by Monte Carlo simulation&lt;br /&gt;
** Simulation algorithms&lt;br /&gt;
** Implementation in R or Python&lt;br /&gt;
** Visualization&lt;br /&gt;
&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
* Random DNA sequences&lt;br /&gt;
* DNA replication&lt;br /&gt;
* DNA transcription&lt;br /&gt;
&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* 6-frame translation&lt;br /&gt;
** Algorithm&lt;br /&gt;
** Visualization&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3: Central Dogma&lt;br /&gt;
&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6744</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6744"/>
		<updated>2026-07-27T07:19:55Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 5 (7/21) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;[https://edu.gitee.com/huntercollege/courses/5124/noticeboard Course link on gitee]&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1, review&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* Linux Tutorial I. Basics. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
* Linux Tutorial II. Advanced&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 4. Linux Intermezzos 1-4&lt;br /&gt;
&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* Linux Tutorial III: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
* Probability by Monte Carlo simulations&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2. Linux commands &amp;amp; BpWraper genomics utilities&lt;br /&gt;
* The Birthday Problem&lt;br /&gt;
** Probability by Monte Carlo simulation&lt;br /&gt;
** Simulation algorithms&lt;br /&gt;
** Implementation in R or Python&lt;br /&gt;
** Visualization&lt;br /&gt;
&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
* Random DNA sequences&lt;br /&gt;
* DNA replication&lt;br /&gt;
* DNA transcription&lt;br /&gt;
&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* 6-frame translation&lt;br /&gt;
** Algorithm&lt;br /&gt;
** Visualization&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3: Central Dogma&lt;br /&gt;
&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6743</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6743"/>
		<updated>2026-07-25T03:25:51Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 10 (8/3) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;[https://edu.gitee.com/huntercollege/courses/5124/noticeboard Course link on gitee]&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1, review&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* Linux Tutorial I. Basics. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
* Linux Tutorial II. Advanced&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* Linux Tutorial III: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
* Probability by Monte Carlo simulations&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2. Linux commands &amp;amp; BpWraper genomics utilities&lt;br /&gt;
* The Birthday Problem&lt;br /&gt;
** Probability by Monte Carlo simulation&lt;br /&gt;
** Simulation algorithms&lt;br /&gt;
** Implementation in R or Python&lt;br /&gt;
** Visualization&lt;br /&gt;
&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
* Random DNA sequences&lt;br /&gt;
* DNA replication&lt;br /&gt;
* DNA transcription&lt;br /&gt;
&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* 6-frame translation&lt;br /&gt;
** Algorithm&lt;br /&gt;
** Visualization&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3: Central Dogma&lt;br /&gt;
&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6742</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6742"/>
		<updated>2026-07-25T03:24:58Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 10 (8/3) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;[https://edu.gitee.com/huntercollege/courses/5124/noticeboard Course link on gitee]&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1, review&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* Linux Tutorial I. Basics. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
* Linux Tutorial II. Advanced&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* Linux Tutorial III: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
* Probability by Monte Carlo simulations&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2. Linux commands &amp;amp; BpWraper genomics utilities&lt;br /&gt;
* The Birthday Problem&lt;br /&gt;
** Probability by Monte Carlo simulation&lt;br /&gt;
** Simulation algorithms&lt;br /&gt;
** Implementation in R or Python&lt;br /&gt;
** Visualization&lt;br /&gt;
&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
* Random DNA sequences&lt;br /&gt;
* DNA replication&lt;br /&gt;
* DNA transcription&lt;br /&gt;
&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* 6-frame translation&lt;br /&gt;
** Algorithm&lt;br /&gt;
** Visualization&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
&amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6741</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6741"/>
		<updated>2026-07-25T03:24:37Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 9 (7/30) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;[https://edu.gitee.com/huntercollege/courses/5124/noticeboard Course link on gitee]&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1, review&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* Linux Tutorial I. Basics. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
* Linux Tutorial II. Advanced&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* Linux Tutorial III: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
* Probability by Monte Carlo simulations&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2. Linux commands &amp;amp; BpWraper genomics utilities&lt;br /&gt;
* The Birthday Problem&lt;br /&gt;
** Probability by Monte Carlo simulation&lt;br /&gt;
** Simulation algorithms&lt;br /&gt;
** Implementation in R or Python&lt;br /&gt;
** Visualization&lt;br /&gt;
&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
* Random DNA sequences&lt;br /&gt;
* DNA replication&lt;br /&gt;
* DNA transcription&lt;br /&gt;
&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* 6-frame translation&lt;br /&gt;
** Algorithm&lt;br /&gt;
** Visualization&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6740</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6740"/>
		<updated>2026-07-25T03:23:56Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 8 (7/28) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;[https://edu.gitee.com/huntercollege/courses/5124/noticeboard Course link on gitee]&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1, review&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* Linux Tutorial I. Basics. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
* Linux Tutorial II. Advanced&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* Linux Tutorial III: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
* Probability by Monte Carlo simulations&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2. Linux commands &amp;amp; BpWraper genomics utilities&lt;br /&gt;
* The Birthday Problem&lt;br /&gt;
** Probability by Monte Carlo simulation&lt;br /&gt;
** Simulation algorithms&lt;br /&gt;
** Implementation in R or Python&lt;br /&gt;
** Visualization&lt;br /&gt;
&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
* Random DNA sequences&lt;br /&gt;
* DNA replication&lt;br /&gt;
* DNA transcription&lt;br /&gt;
&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6739</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6739"/>
		<updated>2026-07-25T03:23:04Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 7 (7/27) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;[https://edu.gitee.com/huntercollege/courses/5124/noticeboard Course link on gitee]&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1, review&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* Linux Tutorial I. Basics. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
* Linux Tutorial II. Advanced&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* Linux Tutorial III: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
* Probability by Monte Carlo simulations&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2. Linux commands &amp;amp; BpWraper genomics utilities&lt;br /&gt;
* The Birthday Problem&lt;br /&gt;
** Probability by Monte Carlo simulation&lt;br /&gt;
** Simulation algorithms&lt;br /&gt;
** Implementation in R or Python&lt;br /&gt;
** Visualization&lt;br /&gt;
&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6738</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6738"/>
		<updated>2026-07-20T05:59:09Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 6 (7/23) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;[https://edu.gitee.com/huntercollege/courses/5124/noticeboard Course link on gitee]&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1, review&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* Linux Tutorial I. Basics. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
* Linux Tutorial II. Advanced&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* Linux Tutorial III: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
* Probability by Monte Carlo simulations&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6737</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6737"/>
		<updated>2026-07-20T05:57:34Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 4 (7/20) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;[https://edu.gitee.com/huntercollege/courses/5124/noticeboard Course link on gitee]&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1, review&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* Linux Tutorial I. Basics. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
* Linux Tutorial II. Advanced&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* Linux Tutorial III: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6736</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6736"/>
		<updated>2026-07-20T05:57:05Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Week 3. Simulations: Central Dogma */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;[https://edu.gitee.com/huntercollege/courses/5124/noticeboard Course link on gitee]&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1, review&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
* Linux Tutorial I. Basics. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
* Linux Tutorial II. Advanced&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* Linux Tutorial III: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6735</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6735"/>
		<updated>2026-07-20T05:56:39Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Week 2. Linux &amp;amp; Command Line Interface */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;[https://edu.gitee.com/huntercollege/courses/5124/noticeboard Course link on gitee]&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1, review&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
* Linux Tutorial I. Basics. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
* Linux Tutorial II. Advanced&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* Linux Tutorial III: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=File:Phylogeny-lecture-slides-2026.pdf&amp;diff=6734</id>
		<title>File:Phylogeny-lecture-slides-2026.pdf</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=File:Phylogeny-lecture-slides-2026.pdf&amp;diff=6734"/>
		<updated>2026-07-19T01:53:38Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: Wikiuser uploaded a new version of File:Phylogeny-lecture-slides-2026.pdf&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6733</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6733"/>
		<updated>2026-07-15T08:18:57Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Learning goals */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;[https://edu.gitee.com/huntercollege/courses/5124/noticeboard Course link on gitee]&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1, review&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
* Linux Tutorial I. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6732</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6732"/>
		<updated>2026-07-15T08:08:11Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 3 (7/16) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;[https://edu.gitee.com/huntercollege/courses/5124/noticeboard Course link on gitee]&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1, review&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
* Linux Tutorial I. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6731</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6731"/>
		<updated>2026-07-15T08:07:26Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;[https://edu.gitee.com/huntercollege/courses/5124/noticeboard Course link on gitee]&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1:&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
* Linux Tutorial I. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6730</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6730"/>
		<updated>2026-07-15T08:06:40Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;centr&amp;gt;[https://edu.gitee.com/huntercollege/courses/5124/noticeboard Course link on gitee]&amp;lt;/centr&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1:&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
* Linux Tutorial I. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6729</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6729"/>
		<updated>2026-07-15T08:05:47Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Week 1. Intro &amp;amp; Phylogenetics */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1:&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
* Linux Tutorial I. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6728</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6728"/>
		<updated>2026-07-15T07:55:05Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Week 1. Intro &amp;amp; Phylogenetics */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
* To prepare for Quiz 1:&lt;br /&gt;
** Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
** Three assignments&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
* Linux Tutorial I. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6727</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6727"/>
		<updated>2026-07-15T07:33:44Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Week 1. Intro &amp;amp; Phylogenetics */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]; Demo: homework submission using Gitee&lt;br /&gt;
* Survey 1: Genomics &amp;amp; Data Science &lt;br /&gt;
* Survey 2: Tree-thinking Skills; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Assignment 3. Parsimony analysis&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
* Linux Tutorial I. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6726</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6726"/>
		<updated>2026-07-15T07:32:29Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 3 (7/16) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]; Demo: homework submission using Gitee&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 1: Genomics &amp;amp; Data Science&amp;lt;/span&amp;gt; &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 2: Tree-thinking Skills&amp;lt;/span&amp;gt;; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
* Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
* Assignment 3. Parsimony analysis&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
* Linux Tutorial I. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6725</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6725"/>
		<updated>2026-07-15T07:31:59Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 3 (7/16) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]; Demo: homework submission using Gitee&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 1: Genomics &amp;amp; Data Science&amp;lt;/span&amp;gt; &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 2: Tree-thinking Skills&amp;lt;/span&amp;gt;; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
* Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* Character reconstruction&lt;br /&gt;
* Phylogenetic analysis: Parsimony analysis&lt;br /&gt;
* Pre-test 2&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
* Linux Tutorial I. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6724</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6724"/>
		<updated>2026-07-15T07:28:30Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 4 (7/20) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]; Demo: homework submission using Gitee&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 1: Genomics &amp;amp; Data Science&amp;lt;/span&amp;gt; &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 2: Tree-thinking Skills&amp;lt;/span&amp;gt;; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
* Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phyogenetics&lt;br /&gt;
* Linux Tutorial I. [[File:Part-1-intro-unix-2026.pdf|thumb]]&lt;br /&gt;
* Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=File:Part-1-intro-unix-2026.pdf&amp;diff=6723</id>
		<title>File:Part-1-intro-unix-2026.pdf</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=File:Part-1-intro-unix-2026.pdf&amp;diff=6723"/>
		<updated>2026-07-15T07:27:35Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6722</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6722"/>
		<updated>2026-07-15T07:26:29Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Course Schedule */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]; Demo: homework submission using Gitee&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 1: Genomics &amp;amp; Data Science&amp;lt;/span&amp;gt; &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 2: Tree-thinking Skills&amp;lt;/span&amp;gt;; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
* Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. Phylogenetics&lt;br /&gt;
* Linux Tutorial I. [[File:Intro-unix-KIZ-2026.pdf|thumb]]&lt;br /&gt;
* Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6721</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6721"/>
		<updated>2026-07-14T02:09:25Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 2 (7/14) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]; Demo: homework submission using Gitee&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 1: Genomics &amp;amp; Data Science&amp;lt;/span&amp;gt; &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 2: Tree-thinking Skills&amp;lt;/span&amp;gt;; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
* Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Install R, Tidyverse, &amp;amp; Positron&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1&lt;br /&gt;
* Linux Tutorial I. [[File:Intro-unix-KIZ-2026.pdf|thumb]]&lt;br /&gt;
* Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=File:Phylogeny-lecture-slides-2026.pdf&amp;diff=6720</id>
		<title>File:Phylogeny-lecture-slides-2026.pdf</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=File:Phylogeny-lecture-slides-2026.pdf&amp;diff=6720"/>
		<updated>2026-07-13T23:51:55Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: Wikiuser uploaded a new version of File:Phylogeny-lecture-slides-2026.pdf&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6719</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6719"/>
		<updated>2026-07-13T23:50:33Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 2 (7/14) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]; Demo: homework submission using Gitee&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 1: Genomics &amp;amp; Data Science&amp;lt;/span&amp;gt; &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 2: Tree-thinking Skills&amp;lt;/span&amp;gt;; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
* Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
conda create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
* Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
* Assignment 2. Tree reroot &amp;amp; gene paralogy&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1&lt;br /&gt;
* Linux Tutorial I. [[File:Intro-unix-KIZ-2026.pdf|thumb]]&lt;br /&gt;
* Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6718</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6718"/>
		<updated>2026-07-13T23:44:44Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Web Links */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): https://docs.posit.co/ide/user/#rstudio-ide-oss-downloads&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]; Demo: homework submission using Gitee&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 1: Genomics &amp;amp; Data Science&amp;lt;/span&amp;gt; &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 2: Tree-thinking Skills&amp;lt;/span&amp;gt;; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
* Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
** Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
** Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1&lt;br /&gt;
* Linux Tutorial I. [[File:Intro-unix-KIZ-2026.pdf|thumb]]&lt;br /&gt;
* Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6717</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6717"/>
		<updated>2026-07-13T23:15:33Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Learning goals */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Tree-thinking&#039;&#039;&#039; skills) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Population-thinking&#039;&#039;&#039; skills) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Algorithm-thinking&#039;&#039;&#039; skills) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Probability-thinking&#039;&#039;&#039; skills) Obtain probabilities by Monte Carlo simulations&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): http://www.rstudio.com/download&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]; Demo: homework submission using Gitee&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 1: Genomics &amp;amp; Data Science&amp;lt;/span&amp;gt; &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 2: Tree-thinking Skills&amp;lt;/span&amp;gt;; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
* Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
** Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
** Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1&lt;br /&gt;
* Linux Tutorial I. [[File:Intro-unix-KIZ-2026.pdf|thumb]]&lt;br /&gt;
* Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6716</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6716"/>
		<updated>2026-07-13T23:12:39Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Course Schedule */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Pop Gen&#039;&#039;&#039;) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Computation&#039;&#039;&#039;) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Phylogenomics&#039;&#039;&#039;) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution &lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): http://www.rstudio.com/download&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1. Intro &amp;amp; Phylogenetics===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]; Demo: homework submission using Gitee&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 1: Genomics &amp;amp; Data Science&amp;lt;/span&amp;gt; &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 2: Tree-thinking Skills&amp;lt;/span&amp;gt;; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
* Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
* Git usages: Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
** Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
** Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1&lt;br /&gt;
* Linux Tutorial I. [[File:Intro-unix-KIZ-2026.pdf|thumb]]&lt;br /&gt;
* Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&lt;br /&gt;
&lt;br /&gt;
===Week 2. Linux &amp;amp; Command Line Interface===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3. Simulations: Central Dogma===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4. Simulations: Population genetics===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5. Microbial genomics===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6715</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6715"/>
		<updated>2026-07-13T14:45:58Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 2 (7/14) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Pop Gen&#039;&#039;&#039;) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Computation&#039;&#039;&#039;) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Phylogenomics&#039;&#039;&#039;) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution &lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): http://www.rstudio.com/download&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]; Demo: homework submission using Gitee&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 1: Genomics &amp;amp; Data Science&amp;lt;/span&amp;gt; &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 2: Tree-thinking Skills&amp;lt;/span&amp;gt;; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
* Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
* Git usages&lt;br /&gt;
** Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Linux Tutorial I. [[File:Intro-unix-KIZ-2026.pdf|thumb]]&lt;br /&gt;
** Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
** Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
** BpWrapper tools&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&amp;lt;/span&amp;gt; Updated sldies: [[File:Intro-unix-KIZ.pdf|thumb]]&lt;br /&gt;
&lt;br /&gt;
===Week 2===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6714</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6714"/>
		<updated>2026-07-13T14:36:42Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 2 (7/14) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Pop Gen&#039;&#039;&#039;) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Computation&#039;&#039;&#039;) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Phylogenomics&#039;&#039;&#039;) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution &lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): http://www.rstudio.com/download&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]; Demo: homework submission using Gitee&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 1: Genomics &amp;amp; Data Science&amp;lt;/span&amp;gt; &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 2: Tree-thinking Skills&amp;lt;/span&amp;gt;; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
* Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Install Linux operating systems (Mac OSX or Windows WSL)&lt;br /&gt;
* Install miniconda &amp;amp; create a new conda environment&lt;br /&gt;
* Git usages&lt;br /&gt;
** Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt;&lt;br /&gt;
* Linux Tutorial I. [[File:Intro-unix-KIZ.pdf|thumb]]&lt;br /&gt;
** Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
** Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
** BpWrapper tools&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&amp;lt;/span&amp;gt; Updated sldies: [[File:Intro-unix-KIZ.pdf|thumb]]&lt;br /&gt;
&lt;br /&gt;
===Week 2===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6713</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6713"/>
		<updated>2026-07-13T14:30:53Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 1 (7/13) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Pop Gen&#039;&#039;&#039;) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Computation&#039;&#039;&#039;) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Phylogenomics&#039;&#039;&#039;) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution &lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): http://www.rstudio.com/download&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]; Demo: homework submission using Gitee&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 1: Genomics &amp;amp; Data Science&amp;lt;/span&amp;gt; &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 2: Tree-thinking Skills&amp;lt;/span&amp;gt;; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
* Assignment 1. Pretest-1. Phylogenetic tree&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Git usages (by Mr Liu)&lt;br /&gt;
** Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt; &lt;br /&gt;
* Linux Tutorial I. [[File:Intro-unix-KIZ.pdf|thumb]]&lt;br /&gt;
** Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
** Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&amp;lt;/span&amp;gt; Updated sldies: [[File:Intro-unix-KIZ.pdf|thumb]]&lt;br /&gt;
&lt;br /&gt;
===Week 2===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=File:Phylogeny-lecture-slides-2026.pdf&amp;diff=6712</id>
		<title>File:Phylogeny-lecture-slides-2026.pdf</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=File:Phylogeny-lecture-slides-2026.pdf&amp;diff=6712"/>
		<updated>2026-07-12T16:22:39Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6711</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6711"/>
		<updated>2026-07-12T16:18:38Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Session 1 (7/13) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Pop Gen&#039;&#039;&#039;) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Computation&#039;&#039;&#039;) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Phylogenomics&#039;&#039;&#039;) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution &lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): http://www.rstudio.com/download&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]; Demo: homework submission using Gitee&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 1: Genomics &amp;amp; Data Science&amp;lt;/span&amp;gt; &lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 2: Tree-thinking Skills&amp;lt;/span&amp;gt;; Handout: Pre-test 1&lt;br /&gt;
* Lecture: Tree-thinking Skills. Lecture slides: [[File:Phylogeny-lecture-slides-2026.pdf]]&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Git usages (by Mr Liu)&lt;br /&gt;
** Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt; &lt;br /&gt;
* Linux Tutorial I. [[File:Intro-unix-KIZ.pdf|thumb]]&lt;br /&gt;
** Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
** Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&amp;lt;/span&amp;gt; Updated sldies: [[File:Intro-unix-KIZ.pdf|thumb]]&lt;br /&gt;
&lt;br /&gt;
===Week 2===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6710</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6710"/>
		<updated>2026-07-12T16:07:28Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Assignments, Quizzes, and a Final Presentatio8n */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Pop Gen&#039;&#039;&#039;) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Computation&#039;&#039;&#039;) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Phylogenomics&#039;&#039;&#039;) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution &lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): http://www.rstudio.com/download&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, in-class quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 14 * 5 pts = 70 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 4 x 30 pts = 120 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 300 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]; Demo: homework submission using Gitee&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 1: Genomics &amp;amp; Data Science&amp;lt;/span&amp;gt; PDF file: [[File:KIZ-survey-1.pdf|thumb]]&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 2: Tree-thinking Skills&amp;lt;/span&amp;gt; PDF file: [[File:Pretest-1.pdf|thumb]]&lt;br /&gt;
* Lecture: Tree-thinking Skills. Lecture slides: [[File:Phylogeny-lecture-slides-2024.pdf]]&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Git usages (by Mr Liu)&lt;br /&gt;
** Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt; &lt;br /&gt;
* Linux Tutorial I. [[File:Intro-unix-KIZ.pdf|thumb]]&lt;br /&gt;
** Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
** Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&amp;lt;/span&amp;gt; Updated sldies: [[File:Intro-unix-KIZ.pdf|thumb]]&lt;br /&gt;
&lt;br /&gt;
===Week 2===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6709</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6709"/>
		<updated>2026-07-12T16:05:46Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Course Schedule */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Pop Gen&#039;&#039;&#039;) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Computation&#039;&#039;&#039;) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Phylogenomics&#039;&#039;&#039;) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution &lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): http://www.rstudio.com/download&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, three (3) quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 12 * 5 pts = 60 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 3 x 30 pts = 60 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 230 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1===&lt;br /&gt;
====Session 1 (7/13)====&lt;br /&gt;
* Introduction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]; Demo: homework submission using Gitee&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 1: Genomics &amp;amp; Data Science&amp;lt;/span&amp;gt; PDF file: [[File:KIZ-survey-1.pdf|thumb]]&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 2: Tree-thinking Skills&amp;lt;/span&amp;gt; PDF file: [[File:Pretest-1.pdf|thumb]]&lt;br /&gt;
* Lecture: Tree-thinking Skills. Lecture slides: [[File:Phylogeny-lecture-slides-2024.pdf]]&lt;br /&gt;
&lt;br /&gt;
====Session 2 (7/14)====&lt;br /&gt;
* Git usages (by Mr Liu)&lt;br /&gt;
** Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt; &lt;br /&gt;
* Linux Tutorial I. [[File:Intro-unix-KIZ.pdf|thumb]]&lt;br /&gt;
** Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
** Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
&lt;br /&gt;
====Session 3 (7/16)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&amp;lt;/span&amp;gt; Updated sldies: [[File:Intro-unix-KIZ.pdf|thumb]]&lt;br /&gt;
&lt;br /&gt;
===Week 2===&lt;br /&gt;
====Session 4 (7/20)====&lt;br /&gt;
====Session 5 (7/21)====&lt;br /&gt;
====Session 6 (7/23)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2&lt;br /&gt;
&lt;br /&gt;
===Week 3===&lt;br /&gt;
====Session 7 (7/27)====&lt;br /&gt;
====Session 8 (7/28)====&lt;br /&gt;
====Session 9 (7/30)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #3&lt;br /&gt;
&lt;br /&gt;
===Week 4===&lt;br /&gt;
====Session 10 (8/3)====&lt;br /&gt;
====Session 11 (8/4)====&lt;br /&gt;
====Session 12 (8/6)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #4&lt;br /&gt;
&lt;br /&gt;
===Week 5===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6708</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6708"/>
		<updated>2026-07-12T15:56:38Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;2-5pm, July 13 - Aug 13, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* (&#039;&#039;&#039;Pop Gen&#039;&#039;&#039;) Achieve an algorithmic understanding of major evolutionary processes (drift, mutation, recombination, and natural selection)&lt;br /&gt;
* (&#039;&#039;&#039;Computation&#039;&#039;&#039;) Implement the evolutionary algorithms in a programming language of choice (R, Python, or others)&lt;br /&gt;
* (&#039;&#039;&#039;Phylogenomics&#039;&#039;&#039;) Compare and analyze genomes in a phylogenetic framework&lt;br /&gt;
* (&#039;&#039;&#039;Microbial evolution&#039;&#039;&#039;) Understand the distinctions between microbial and eukaryotic evolution &lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): http://www.rstudio.com/download&lt;br /&gt;
* R/Tidyverse reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* A textbook on microbial genomics &amp;amp; evolution by the instructor: [https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 &#039;&#039;Digital Genomes (Qiu 2026)&#039;&#039;]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and a Final Presentatio8n==&lt;br /&gt;
Student performance will be evaluated by attendance, three (3) quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 12 * 5 pts = 60 pts&lt;br /&gt;
* Daily assignments:  8 x 10 pts  = 80 pts&lt;br /&gt;
* Open-Book Quizzes: 3 x 30 pts = 60 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 230 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1===&lt;br /&gt;
====Session 1, 7/13 (Monday)====&lt;br /&gt;
* Introudction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Gitee accounts &amp;amp; usage [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]; Demo: homework submission using Gitee&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 1: Genomics &amp;amp; Data Science&amp;lt;/span&amp;gt; PDF file: [[File:KIZ-survey-1.pdf|thumb]]&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 2: Tree-thinking Skills&amp;lt;/span&amp;gt; PDF file: [[File:Pretest-1.pdf|thumb]]&lt;br /&gt;
* Lecture: Tree-thinking Skills. Lecture slides: [[File:Phylogeny-lecture-slides-2024.pdf]]&lt;br /&gt;
&lt;br /&gt;
====Session 2, 7/14 (Tuesday)====&lt;br /&gt;
* Git usages (by Mr Liu)&lt;br /&gt;
** Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt; &lt;br /&gt;
* Linux Tutorial I. [[File:Intro-unix-KIZ.pdf|thumb]]&lt;br /&gt;
** Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
** Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
&lt;br /&gt;
====Session 3, 7/16 (Thursday)====&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. tree terms &amp;amp; tree manipulations&amp;lt;/span&amp;gt; (20 pts); open-book&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&amp;lt;/span&amp;gt; Updated sldies: [[File:Intro-unix-KIZ.pdf|thumb]]&lt;br /&gt;
&lt;br /&gt;
===Week 2===&lt;br /&gt;
===Week 3===&lt;br /&gt;
===Week 4===&lt;br /&gt;
====Session 13 (8/10)====&lt;br /&gt;
====Session 14 (8/11)====&lt;br /&gt;
* Presentations&lt;br /&gt;
====Session 15 (8/13)====&lt;br /&gt;
* Presentations&lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Monte_Carlo_Club&amp;diff=6707</id>
		<title>Monte Carlo Club</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Monte_Carlo_Club&amp;diff=6707"/>
		<updated>2026-07-03T19:49:55Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Simulations: Non-adaptive selection */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;__FORCETOC__&lt;br /&gt;
=Mcirobiome &amp;amp; tick pop gen: Ecology &amp;amp; Evolution perspectives (Spring &amp;amp; Summer 2026)=&lt;br /&gt;
==Simulations: Non-adaptive selection==&lt;br /&gt;
* Heterosis: maintenance of genomic diversity by heterozygote advantages&lt;br /&gt;
* Stability of the food web. In microbial system: https://www.nature.com/articles/s41467-018-05308-z&lt;br /&gt;
* May threshold: Ecological stability: simulations by Brian : https://github.com/aclaAPCSA/ACLAatQiuLab/blob/main/mayFunction.ipynb&lt;br /&gt;
* Prey-predator enrichment&lt;br /&gt;
&lt;br /&gt;
==Simulations: artificial immune systems (AIS)==&lt;br /&gt;
* r-contiguous symbols rule for antigen recognition: https://www.pnas.org/doi/10.1073/pnas.90.5.1691&lt;br /&gt;
* Negative selection: weed out self-antigens, http://dpnm.postech.ac.kr/ANM/resources/papers/09_Nature/hofmeyr_forrest.pdf&lt;br /&gt;
* Clonal selection algorithm for learning: https://ieeexplore.ieee.org/abstract/document/1011539&lt;br /&gt;
&lt;br /&gt;
==Pan-genome graphs:  an efficient computational representation of whole-genomic variations==&lt;br /&gt;
* A summary (including data standards &amp;amp; libraries): https://pangenome.github.io/&lt;br /&gt;
* A review paper: https://www.nature.com/articles/s41588-024-02029-6&lt;br /&gt;
* A paper (Huamn &amp;amp; Drosophila models): https://www.nature.com/articles/s41587-023-01793-w&lt;br /&gt;
* SARS-CoV-2 application: whole-genome alignment with Twilight &amp;amp; compression by [https://github.com/TurakhiaLab/panman panMAN]: https://www.amd.com/en/developer/resources/technical-articles/2026/ucsd-researchers-use-amd-gpus-to-build-the-worlds-large.html&lt;br /&gt;
* An application: PHI, for pangenome-based genotyping: https://github.com/at-cg/PHI&lt;br /&gt;
* Build bacterial pangenome with Panaroo: https://link.springer.com/article/10.1186/s13059-020-02090-4&lt;br /&gt;
* Phylogenetic compression of bacterial genome data: MiniPhy &amp;amp; PanMan algorithms; https://www.nature.com/articles/s41579-026-01303-0&lt;br /&gt;
&lt;br /&gt;
==Tick genomics &amp;amp; native immunity==&lt;br /&gt;
===Papers found by Hui===&lt;br /&gt;
* Abbas, M. N., Jmel, M. A., Mekki, I., Dijkgraaf, I., &amp;amp; Kotsyfakis, M. (2023). Recent Advances in Tick Antigen Discovery and Anti-Tick Vaccine Development. International journal of molecular sciences, 24(5), 4969. https://doi.org/10.3390/ijms24054969. I found this review article to be the most complete and well-rounded one. It summarizes recent progress in anti-tick vaccine development, especially antigen-based strategies that target tick feeding, reproduction, attachment, and pathogen transmission. The authors also explain how modern approaches such as genomics, transcriptomics, proteomics, and reverse vaccinology have helped identify new vaccine candidates from tick tissues including the salivary glands, midgut, eggs, Malpighian tubules, and cement. The review covers several major antigens and vaccine targets, including Bm86/Bm95, ferritin, aquaporins, subolesin, Salp15, metalloproteases, serpins, and the recent 19ISP mRNA vaccine targeting Ixodes scapularis salivary proteins.&lt;br /&gt;
* Allen J.R., Humphreys S.J. Immunisation of guinea pigs and cattle against ticks. Nature. 1979;280:491–493. doi: 10.1038/280491a0. https://www-nature-com.proxy.wexler.hunter.cuny.edu/articles/280491a0. The first review article referred to this early study, so I thought it was important to include. This paper is an early proof-of-concept showing that hosts can be immunized against ticks using antigens extracted from the internal organs of partially fed ticks, especially the midgut and reproductive organs. It helped establish the idea that internal tick antigens, not only salivary proteins, could be useful vaccine targets.&lt;br /&gt;
* Hart, T. M., Cui, Y., Telford, S. R., Marín-López, A., Calloway, K., Dai, Y., Matias, J., DePonte, K., Jaycox, J., DeBlasio, M., Hoornstra, D., Belperron, A. A., Cibichakravarthy, B., Johnson, E. E., Alameh, M. G., Dwivedi, G., Hovius, J. W. R., Bockenstedt, L. K., Weissman, D., Ring, A. M., … Fikrig, E. (2025). Tick feeding or vaccination with tick antigens elicits immunity to the Ixodes scapularis exoproteome in guinea pigs and humans. Science translational medicine, 17(791), eads9207. https://doi.org/10.1126/scitranslmed.ads9207. This paper developed a high-throughput yeast display platform called IscREAM, which can screen antibody responses against more than 3,000 predicted extracellular proteins from Ixodes scapularis. Using vaccinated guinea pigs, Lyme disease patients, and a naturally tick-resistant individual, the authors identified 199 immunogenic tick antigens. They also showed that both salivary and cement proteins may contribute to acquired tick resistance.&lt;br /&gt;
* Narasimhan, Sukanya et al. “Immunity against Ixodes scapularis salivary proteins expressed within 24 hours of attachment thwarts tick feeding and impairs Borrelia transmission.” PloS one vol. 2,5 e451. 16 May. 2007, https://journals.plos.org/plosone/article?id=10.1371%2Fjournal.pone.0000451. This article focuses on early Ixodes scapularis salivary proteins that are important for tick feeding and Borrelia burgdorferi transmission. The authors show that the tick salivary gland transcriptome and proteome change during feeding, especially between the first 24 hours of attachment and later feeding stages. Importantly, they found that guinea pigs repeatedly exposed only to the first 24 hours of tick feeding developed acquired tick immunity, shown by inflammation at the bite site, rapid tick rejection, reduced engorgement, and impaired Borrelia transmission.&lt;br /&gt;
&lt;br /&gt;
===Tick genomics===&lt;br /&gt;
* Mini-review: [https://www.frontiersin.org/journals/cellular-and-infection-microbiology/articles/10.3389/fcimb.2018.00176/full Kitsou &amp;amp; Pal (2018).] Ixodes immune responses against Lyme disease pathogens. Font Cell Infect Microb.&lt;br /&gt;
* The IMD pathway: [https://journals.asm.org/doi/10.1128/mbio.00703-22 Sidak-Loftis et al (2022)]; mBio&lt;br /&gt;
* The PERK pathway: [https://journals.asm.org/doi/10.1128/msphere.00321-23 Rosche et al (2023)]; mSphere&lt;br /&gt;
* Tick microbiome: [https://journals.asm.org/doi/10.1128/spectrum.01404-23 Paulson et al (2023);] Microbio Spectrum&lt;br /&gt;
* The tick genome paper: [https://www.nature.com/articles/s41588-022-01275-w De et al (2023).] Nature Genetics&lt;br /&gt;
* Tcik genome resource:&lt;br /&gt;
&lt;br /&gt;
==Microbiome &amp;amp; Evolution of Complexity: the Holobionts/Gia hypothesis vs Non-adaptive selection==&lt;br /&gt;
* [https://press.uchicago.edu/ucp/books/book/chicago/N/bo244157692.html Non-adaptive Selection], by John Damuth &amp;amp; Lev Ginzburg&lt;br /&gt;
* C.H. Lean, W.F. Doolittle, &amp;amp; J.P. Bielawski, Community-level evolutionary processes: Linking community genetics with &#039;&#039;&#039;replicator-interactor theory&#039;&#039;&#039;, Proc. Natl. Acad. Sci. U.S.A. 119 (46) e2202538119, https://doi.org/10.1073/pnas.2202538119 (2022).&lt;br /&gt;
* Guilhem Doulcier, Amaury Lambert, Silvia De Monte, Paul B Rainey (2020) Eco-evolutionary dynamics of nested Darwinian populations and the emergence of community-level heredity eLife 9:e53433. https://doi.org/10.7554/eLife.53433 &lt;br /&gt;
* Daniel W. McShea, Steve C. Wang, Robert N. Brandon, A quantitative formulation of biology&#039;s first law, Evolution, Volume 73, Issue 6, 1 June 2019, Pages 1101–1115, https://doi.org/10.1111/evo.13735 (the &amp;quot;&#039;&#039;&#039;The zero-force evolutionary law (ZFEL)&#039;&#039;&#039;&amp;quot;)&lt;br /&gt;
* J. Madhusoodanan, Do hosts and their microbes evolve as a unit?, Proc. Natl. Acad. Sci. U.S.A. 116 (29) 14391-14394, https://doi.org/10.1073/pnas.1908139116 (2019).&lt;br /&gt;
* Douglas AE, Werren JH. 2016.Holes in the Hologenome: Why Host-Microbe Symbioses Are Not Holobionts. mBio7:10.1128/mbio.02099-15.https://doi.org/10.1128/mbio.02099-15&lt;br /&gt;
* Ilana Zilber-Rosenberg, Eugene Rosenberg, Role of microorganisms in the evolution of animals and plants: &#039;&#039;&#039;the hologenome theory&#039;&#039;&#039; of evolution, FEMS Microbiology Reviews, Volume 32, Issue 5, August 2008, Pages 723–735, https://doi.org/10.1111/j.1574-6976.2008.00123.x&lt;br /&gt;
* Seth R Bordenstein &amp;amp; Kevin R Theis (2015). Host Biology in Light of the Microbiome: Ten Principles of Holobionts and Hologenomes. PLoS Biology. https://doi.org/10.1371/journal.pbio.1002226&lt;br /&gt;
&lt;br /&gt;
=Coevolution, Misinformation, and Sex (Summer &amp;amp; Fall 2025)=&lt;br /&gt;
==Schedule &amp;amp; Notes==&lt;br /&gt;
* Orientation: 6/19/2025. Brian, Ryan, Emma, Sebastian, Will&lt;br /&gt;
* Week 1. (June 22, 2025, Monday). &lt;br /&gt;
** Presentation: Network evolution and RNA world (by Emma)&lt;br /&gt;
** Presentation: Review last season&#039;s Python work (by Sebastian)&lt;br /&gt;
** Presentation: Python Automata class (by Will)&lt;br /&gt;
** Presentation: Automata &amp;amp; biological applications (by Brandon)&lt;br /&gt;
** Presentation: k-ras data analysis, normalized to DMSO with fold change; linear model; clustering by heatmap (Ahava)&lt;br /&gt;
* Week 2. (June 30, 2025, Monday) &lt;br /&gt;
** Sebastian: Shannon entropy. To do: simulate the small-sample size effect (reduction of uncertainty)&lt;br /&gt;
** Ryan: genetic code &amp;amp; 6-frame translation. To do: 6-frame translation of ospC sequence&lt;br /&gt;
** Brian and Will: Miller Appendix F, Test 1. Adaptation to a fixed string (with genetic algorithm). To do: Test 2. pattern recognition&lt;br /&gt;
** Brandon: k-ras data analysis, gene set enrichment analysis. To do: pick random genes from clusters to validate the results of linear models&lt;br /&gt;
* Week 3. (July 7, 2025)&lt;br /&gt;
** Ryan: done with 6-frame translation &amp;amp; visualization code&lt;br /&gt;
** Sebastian: Simulated entropy given sample size, showing the loss of uncertainty with small sample sizes&lt;br /&gt;
** Brian: done with Miller Appendix F2, evolving automata for pattern recognition&lt;br /&gt;
** Brandon, Ahava, and Lenny: k-ras data heatmaps, preliminary GSEA results, and interaction plots&lt;br /&gt;
* Week 4 (July 14, 2025)&lt;br /&gt;
** Ryan and Sebastian: Mutual info implemented&lt;br /&gt;
** Brian: Implement F3, co-evolving problem set (to avoid over-fitting)&lt;br /&gt;
** Brandon: an automaton model of two-host sensing; test generalist vs specialist; test cost of complexity&lt;br /&gt;
* Week 5 (July 21, 2025)&lt;br /&gt;
** Ryan and Sebastian: Simulate correlated counts; constructed function for mutual information (MI) and D-print&lt;br /&gt;
** Brandon &amp;amp; Brian: evolve two-host sensing; implement tradeoff under the constraint of finite number of states&lt;br /&gt;
* Week 6 (Tuesday, July 29, 2025)&lt;br /&gt;
** Ryan &amp;amp; Sebastian: both finished functions for MIR and bootstrap test&lt;br /&gt;
** Next task: implement relative conditional entropy (RCE) as an asymmetric measure of information content between two loci&lt;br /&gt;
* Week 7 (break)&lt;br /&gt;
* Week 8 (Wed, Aug 13, 2025)&lt;br /&gt;
** Ryan &amp;amp; Sebastian: RCE on Lyme data with bootstrap tests; Quantify genetic differentiation between two populations: Fst (hz, pi) &amp;amp; Jensen-Shannon divergence&lt;br /&gt;
** Brian: automata model for generalists and specialist&lt;br /&gt;
** Brandon: hybrid models with two loci, one under immune selection, another under adaptive selection. Results similar to immune-selection model alone&lt;br /&gt;
* Week 9 (Friday, Aug 21, 2025)&lt;br /&gt;
** End-of-summer party&lt;br /&gt;
** Bootstrap on RCE, by break linkage&lt;br /&gt;
** Agricultural optimization problem&lt;br /&gt;
** Quarto book/website to summarize all work (code and visualization)&lt;br /&gt;
&lt;br /&gt;
== Reading list==&lt;br /&gt;
* &#039;&#039;&#039;Coevolutionary automata&#039;&#039;&#039; John Miller (2023). [https://www.sfipress.org/books/ex-machina Ex Machina]. [https://github.com/SantaFeInstitute/ExMachina?tab=readme-ov-file PDF book on GitHub]&lt;br /&gt;
* &#039;&#039;&#039;Network evolution&#039;&#039;&#039; (automata model): Phylip Ball, Nautilus (2025): [https://nautil.us/the-strange-inevitability-of-evolution-235189/ The Strange Inevitability of Evolution] &lt;br /&gt;
* David P. Barash (2025). [https://nautil.us/evolution-is-really-not-that-into-sex-237397/ Evolution Is Really Not That Into Sex] &lt;br /&gt;
* Hartfield &amp;amp; Keighterley (2012). &amp;quot;Current hypotheses for the evolution of sex and recombination&amp;quot;.https://pubmed.ncbi.nlm.nih.gov/22691203/&lt;br /&gt;
* wikipedia chapter on mutual information: https://en.wikipedia.org/wiki/Mutual_information&lt;br /&gt;
* David MacKay (2004). Chapter 19 &amp;quot;Why have Sex? Information Acquisition and Evolution&amp;quot;&lt;br /&gt;
* Chanda et al (2020). &amp;quot;Information Theory in Computational Biology: Where We Stand Today.&amp;quot; Entropy 2020, 22, 627. https://doi.org/10.3390/e22060627&lt;br /&gt;
&lt;br /&gt;
==Goals and Applications==&lt;br /&gt;
* network evolution and automata model of evolution (e.g., two-host adaptation of Lyme pathogens)&lt;br /&gt;
* genome-wide linkage: multivariate/total correlation&lt;br /&gt;
* genetic costs &amp;amp; benefits of recombination&lt;br /&gt;
&lt;br /&gt;
==Tools &amp;amp; Tasks==&lt;br /&gt;
* Replicate/Implement simulations in MacKay (204) Chapter 19 &lt;br /&gt;
* Python package &amp;lt;code&amp;gt;dit&amp;lt;/code&amp;gt; (discrete info theory): https://dit.readthedocs.io/en/latest/measures/shannon.html#mutual-information&lt;br /&gt;
* R packages for information theory:&lt;br /&gt;
** Distances: philentropy, https://github.com/drostlab/philentropy; https://cran.r-project.org/web/packages/philentropy/vignettes/Information_Theory.html&lt;br /&gt;
** Package: entropy&lt;br /&gt;
** Package: infotheo&lt;br /&gt;
&lt;br /&gt;
==Applications==&lt;br /&gt;
* genome-wide linkage: multivariate/total correlation&lt;br /&gt;
&lt;br /&gt;
=Bb pangenomics project (Fall 2024, with Brian &amp;amp; students)=&lt;br /&gt;
==Reading list==&lt;br /&gt;
* Pop-gen Theory: Baumdicker &amp;amp; Kupczok (2023), &amp;quot;Tackling the Pangenome Dilemma Requires the Concerted Analysis of Multiple Population Genetic Processes&amp;quot;, GBE. https://academic.oup.com/gbe/article/15/5/evad067/7137407?login=false &lt;br /&gt;
* Akther et al (2024). &amp;quot;Natural selection and recombination at host-interacting lipoprotein loci drive genome diversification of Lyme disease and related bacteria&amp;quot; [https://journals.asm.org/doi/10.1128/mbio.01749-24 mBio 0:e01749-24.]&lt;br /&gt;
* Tonkin-Hill  et al (2023). &amp;quot;Robust analysis of prokaryotic pangenome gene gain and loss rates with Panstripe&amp;quot;. [https://genome.cshlp.org/content/33/1/129.full?sid=07484132-d2b7-4fa7-bde0-8d182c64929b Genome Research]&lt;br /&gt;
* Gautreau et al (2020). &amp;quot;PPanGGOLiN: Depicting microbial diversity via a partitioned pangenome graph&amp;quot;. [https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1007732 PLoS Comp Bio.]&lt;br /&gt;
&lt;br /&gt;
==Project Description==&lt;br /&gt;
* Project Goals&lt;br /&gt;
# Species-specific gene gains &amp;amp; losses&lt;br /&gt;
# Correlation with phenotypes (e.g., geography, tick &amp;amp; host association, virulence &amp;amp; pathogenicity)&lt;br /&gt;
&lt;br /&gt;
* Project outline&lt;br /&gt;
# Phase I. Run ppangglin pipeline for 78 Bb main chromosomes: obtain gene presence/absence&lt;br /&gt;
# Phase II. Run phylogenetic reconstruction: obtain branch-specific gene gains and losses&lt;br /&gt;
# Phase III. Run linear regression: obtain gene accumulation curves&lt;br /&gt;
# Phase IV. Plasmid-borne genes&lt;br /&gt;
&lt;br /&gt;
==Preliminary results==&lt;br /&gt;
* PPanGGoLiN output 1: [https://borreliabase.org/~wgqiu/Ushaped_plot.html U-shaped plot for 47 Bb genomes]&lt;br /&gt;
* PPanGGoLiN output 2: [https://borreliabase.org/~wgqiu/tile_plot.html Heatmap of the presence &amp;amp; absence of genes]&lt;br /&gt;
&lt;br /&gt;
=Season VII. Spring &amp;amp; Summer 2024=&lt;br /&gt;
==Monte Carlo Summer Coding Camp==&lt;br /&gt;
* Github repository: https://github.com/weigangq/MonteCarlo&lt;br /&gt;
* For each problem,&lt;br /&gt;
# Step 1. Students will first learn the use of simulator functions (in R and Python)&lt;br /&gt;
# Step 2. Students will write out pseudo code &lt;br /&gt;
# Step 3. Students will implement in R and Python&lt;br /&gt;
# Step 4. Students will make a summary plot&lt;br /&gt;
# Code will be posted on a github as a permanent citable record&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Caption text&lt;br /&gt;
|-&lt;br /&gt;
! Problem !! Simulator !! Pseudo-code !! Data Structure || Expected figure/output !! Learning goal&lt;br /&gt;
|-&lt;br /&gt;
| Birthday: what is the probability of having at least one shared birthday in a class of N students? || uniform samplers. R: &amp;lt;code&amp;gt;sample(); Python: &amp;lt;code&amp;gt;numpy.random.uniform()  || ?? || list/vector || XY plot: X-axis: class size; Y-axis: prob of at least one shared birthday || simulation as experiments; sampling w/ and w/o replacements &lt;br /&gt;
|-&lt;br /&gt;
| Gene frequency: Frequentist approach: maximum likelihood (MLE), Prob(k given p &amp;amp; N): What is the sampling error of a gene frequency given that k alleles are found in a sample of N haploid individuals? || binomial sampler || ?? || ? || mean &amp;amp; variance/standard error || sampling error; maximum likelihood&lt;br /&gt;
|-&lt;br /&gt;
| Gene frequency: Bayesian approach, Prob(p given k &amp;amp; N): what is the gene frequency given k alleles are found in a sample of N haploid individuals || beta sampler || ?? || ? || distribution of posterior probs given success=1,10,100 and failure=9,90,900  || prior &amp;amp; posterior probabilities; Bayes rule; updating rules for beta function  &lt;br /&gt;
|-&lt;br /&gt;
| Monte Hall problem: Should you switch your pick after the host reveals one of the three doors, behind only one of which is a car? || Uniform sampler || ?? || ?? || ?? || conditional/Bayesian prob by simulation&lt;br /&gt;
|-&lt;br /&gt;
| Blending inheritance: why would evolution grind to a halt if genes blend and children are &amp;quot;averages&amp;quot; of two parents? || Uniform sampler || input: one parental populations of N=1000 individuals, each with equal probability of 0&#039;s and 1&#039;s;  a child is produced by randomly selecting a mom and dad and averaging their values; repeat for G generations. Output: genetic diversity in each generation (sd) || ?? || ?? || genetic diversity&lt;br /&gt;
|-&lt;br /&gt;
| Mendenlian inheritance: How would genetic diversity change with random mating?  || Uniform sampler || input: one parental (diploid) populations of N=1000, with two homozygotes &amp;quot;AA&amp;quot; (e.g., 900) and &amp;quot;GG&amp;quot; (e.g., 100);  a child is produced by randomly selecting a mom and dad, followed by another random selection of gametes (A or G from each parent).  Repeat for G generations. Output: genetic diversity in each generation || ?? || ?? || heterozygosity = 1 - sum(p^2)&lt;br /&gt;
|-&lt;br /&gt;
| Genetic drift  || Binomial sampler || Start with a haploid population of 100 individuals with &amp;quot;A&amp;quot; and another 100 individual with &amp;quot;G&amp;quot;. Reproduce for 500 generations while maintaining the total population size of N=200. Track the frequency of the &amp;quot;A&amp;quot; allele, Repeat by using 10X number of individuals || ?? || ?? || &lt;br /&gt;
|-&lt;br /&gt;
| Simulate a genome  || Uniform sampler || The SARS-CoV-2 genome has ~29K bases that is 70% AT. Simulate the genome with the same AT% || ?? || ?? || &lt;br /&gt;
|-&lt;br /&gt;
| Central Dogma  || Uniform sampler || Simulate DNA duplication, Transcription, and 6-frame translation || ?? || ?? || &lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Summer Projects &amp;amp; Readings==&lt;br /&gt;
* Borrelia evolution (Brandon, Saymon, Brian, Che)&lt;br /&gt;
** Nanopore sequencing, evolutionary analysis&lt;br /&gt;
** Co-evolution simulations &amp;amp; modeling (Brandon)&lt;br /&gt;
** ospC multi-epitope model &lt;br /&gt;
* BorreliaBase (Lia &amp;amp; Yozen): &lt;br /&gt;
** Shiny app deployment (done)&lt;br /&gt;
** code refactoring; &lt;br /&gt;
** variable-plasmid new genome data&lt;br /&gt;
** new visualizations (treemap; collapsable trees); &lt;br /&gt;
** gtf files for downloading&lt;br /&gt;
* Miscellaneous collaborative questions&lt;br /&gt;
** HIV intra-host evolution (Dr Evering)&lt;br /&gt;
** scRNA (Dr Melendaz)&lt;br /&gt;
* Readings: Estimating gene frequencies &amp;amp; testing natural selection in bacterial and viral samples&lt;br /&gt;
** [https://journals.asm.org/doi/10.1128/iai.00244-23?url_ver=Z39.88-2003#supplementary-materials Schifflett et al (2023). OpsC alleles in mammals]. To do: Download data set and replicate the figures (including diversity indices) (Laziz)&lt;br /&gt;
** [https://academic.oup.com/pnasnexus/article/2/8/pgad234/7236418#419492093 Combs et al (2023). OspC frequencies in birds and other hosts.] To do: Supplemental Table S1 &amp;amp; S3&lt;br /&gt;
** [https://elifesciences.org/articles/35962 McCrone et al. Influenza within-host variants]. To do: Download &amp;amp; run model code from github. The presence/absence model, compare with the previous tick model (see below)&lt;br /&gt;
** A method for estimating allele frequencies in bacterial samples: [https://academic.oup.com/genetics/article/158/3/1383/6049706 Anderson &amp;amp; Scheet (2001).]&lt;br /&gt;
** [https://www.mdpi.com/1999-4915/13/7/1216 Zhu et al (2021). A model to estimate viral fitness.] To do: implement the github code&lt;br /&gt;
&lt;br /&gt;
==Spring schedule==&lt;br /&gt;
* Week 2 (Feb 1)&lt;br /&gt;
** Presentation: Wayne (HIV tropism). Data: gp120 seqs. Code: from Arezoo&lt;br /&gt;
** Presentation: Aura &amp;amp; Michelle (Single-cell sequencing on endometrial carcinoma&lt;br /&gt;
** Presentation: Susmita &amp;amp; Lily (Borrelia antigens). Data: variability analysis of chosen antigens&lt;br /&gt;
&lt;br /&gt;
* Week 1 (Jan 25)&lt;br /&gt;
** Presentation: Esther (GBS by GEMMA)&lt;br /&gt;
** Project update: Borrelia transcriptomics visualization &amp;amp; Shiny App&lt;br /&gt;
** Presentation: Liann (Borrelia replication &amp;amp; polyploidy)&lt;br /&gt;
** Project update: Arezoo (dummy data set II, phylogenetic random)&lt;br /&gt;
** An R package to read Prism data: [https://cran.r-project.org/web/packages/pzfx/vignettes/pzfx.html The pzfx package]&lt;br /&gt;
&lt;br /&gt;
==Projects &amp;amp; Readings==&lt;br /&gt;
* Borreliella genomics: &lt;br /&gt;
** Genome &amp;amp; gene variability: [http://test.borreliabase.org/ BorreliaBase]&lt;br /&gt;
** Transcriptomes by [https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0164165 Arnold et al (2016). PLoS One]&lt;br /&gt;
** Diagnostic antigens by [https://jcm.asm.org/content/57/12/e01142-19 Arumugam et al (2019). J. Clin. Micro.]&lt;br /&gt;
** Plasmid partitioning by [https://www.nature.com/articles/s41467-022-34876-4 Takacs et al (2022). Nat Communications]&lt;br /&gt;
** DbpA variability in European isolates, by [https://pubmed.ncbi.nlm.nih.gov/16530482/ Schulte-Spechtel et al (2006). Int J. Med Microbiol]&lt;br /&gt;
** Non-abundant antigens by [https://journals.asm.org/doi/10.1128/msystems.01087-23?utm_source=literatum&amp;amp;utm_medium=email&amp;amp;utm_campaign=alert-emails Yaş et al (2023). mSystems]&lt;br /&gt;
* Single-cell transcriptomes of cancer cells&lt;br /&gt;
** [https://www.nature.com/articles/s41467-022-33982-7 Ren et al (2022). Nat Comm.] &lt;br /&gt;
** [https://pubmed.ncbi.nlm.nih.gov/35560676/ Yu et al (2022). Cell Prolif]&lt;br /&gt;
** [https://satijalab.org/seurat/articles/pbmc3k_tutorial Tutorial: basic pipeline with Suerat]&lt;br /&gt;
** [https://satijalab.org/seurat/articles/pbmc3k_tutorial Tutorial: annotation with singleR]&lt;br /&gt;
* Protein classifiers&lt;br /&gt;
** Simulated data&lt;br /&gt;
** HIV tropism&lt;br /&gt;
*** Salmi et al (2013). J. Leukocyte Biology. [https://academic.oup.com/jleukbio/article-abstract/93/1/113/6936163?redirectedFrom=fulltext&amp;amp;login=false Macrophage-tropic HIV-1 variants from brain demonstrate alterations in the way gp120 engages both CD4 and CCR5]&lt;br /&gt;
*** Sequences from Gray et al (2011). Retrovirology. [https://retrovirology.biomedcentral.com/articles/10.1186/1742-4690-8-42 Conformational alterations in the CD4 binding cavity of HIV-1 gp120 influencing gp120-CD4 interactions and fusogenicity of HIV-1 envelopes derived from brain and other tissues]&lt;br /&gt;
** Toxin specifications&lt;br /&gt;
* Co-evolution: simulations &amp;amp; applications&lt;br /&gt;
** Co-evolutionary antomata: by John Miller (2022). [https://github.com/SantaFeInstitute/ExMachina Github]&lt;br /&gt;
** Evolving synthetic proteins by [https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10403280/ Yang et al (2023) Science]&lt;br /&gt;
&lt;br /&gt;
=Season VI. Language Models, Automata, and Evolutionary Games (Spring, Summer &amp;amp; Fall 2023)=&lt;br /&gt;
==Week 2. Lab meeting (June 8, 2023)==&lt;br /&gt;
* GBS microbial genomics. Participants: Eamen, Esther&lt;br /&gt;
* Conotoxin project: unsupervised clustering by Snekmer (amino acid recoding + kmer). Participants: Arezoo, Jessy&lt;br /&gt;
* Red Queen evolution: Reinforcement learning implementation of tic-tac-toe: [https://github.com/ShangtongZhang/reinforcement-learning-an-introduction RL codes]. Participants: Brandon&lt;br /&gt;
* Card game: [https://evolution-outreach.biomedcentral.com/articles/10.1186/s12052-015-0039-2 Gibson et al (2015). &amp;quot;The Red Queen’s Race: An Experimental Card Game to Teach Coevolution&amp;quot;]. Participants: all&lt;br /&gt;
==Week 1. Summer kickoff at Rockefeller University (June 1, 2023)==&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;1&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File:Summer-kickoff-RU-2023.jpg&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
==Projects &amp;amp; Papers==&lt;br /&gt;
* Red Queen (RH) evolution through co-evolutionary automata models (CAM model)&lt;br /&gt;
** Github &amp;amp; Free PDF book: https://github.com/SantaFeInstitute/ExMachina&lt;br /&gt;
** With Reinforcement learning (RL) &amp;amp; finite Markov decision processes (MDPs): [http://incompleteideas.net/book/the-book-2nd.html Book in PDF]; Python implementations: [https://github.com/ShangtongZhang/reinforcement-learning-an-introduction Python codes] &lt;br /&gt;
* Identification of toxin specificity through machine learning and protein language models&lt;br /&gt;
** Readings &amp;amp; Github repository: [https://github.com/weigangq/cono-models/tree/main https://]https://github.com/weigangq/cono-models&lt;br /&gt;
* Design of vaccinogen through polyclonal data analysis &amp;amp; antigenic landscape&lt;br /&gt;
** Lyme antigen design with genetic algorithms (GA): https://www.nature.com/articles/s41396-021-01089-4&lt;br /&gt;
** Mutational landscape using MSA protein language models: https://www.sciencedirect.com/science/article/pii/S2211124722018526&lt;br /&gt;
* Generative art contest, hosted by the &#039;&#039;Chance&#039;&#039; magazine of ASA (Am Stat Association)&lt;br /&gt;
** Contest website (requirements, categories, and deadline: Aug 31, 2023): https://community.amstat.org/chancemagazinecontest/home&lt;br /&gt;
** Software tool 1. R package &amp;quot;generativeart&amp;quot;: https://github.com/cutterkom/generativeart&lt;br /&gt;
** Software tool 2. R package &amp;quot;GenerativeR&amp;quot;, [https://github.com/MartinMSPedersen/GenerativeR generative art with R and Shiny] &lt;br /&gt;
** Software tool 3. R package &amp;quot;aRtsy&amp;quot;, [https://github.com/koenderks/aRtsy Generative Art with ggplot2]&lt;br /&gt;
* Reinforcement learning (RL) &amp;amp; finite Markov Decision Processes (MDPs) as a adaptive learning model&lt;br /&gt;
** Code re-implemented in Python: https://github.com/ShangtongZhang/reinforcement-learning-an-introduction&lt;br /&gt;
** Book PDF: http://incompleteideas.net/book/the-book-2nd.html&lt;br /&gt;
&lt;br /&gt;
=Season V. Genes, Memes, and Machines (Spring, Summer &amp;amp; fall 2022)=&lt;br /&gt;
* A journal club to continue the exploration of the link between evolution &amp;amp; learning. &lt;br /&gt;
==Week 1==&lt;br /&gt;
* A primer of information, presented by Winston [[File:Information Theory Primer-Winston.pdf|thumbnail]]&lt;br /&gt;
* ITSNTS theory of selection unit, presented by Anh. [[File:ITSNTS - Anh Pham.pdf|thumbnail]]&lt;br /&gt;
==Week 2==&lt;br /&gt;
* Information theory of Individuality, presented by Anh [[File:Information_theory_of_individual_Anh_Pham.pdf]]&lt;br /&gt;
* Information definition of liveliness, presented by Winston [[File:Aliveness.pdf]]&lt;br /&gt;
==Week 3==&lt;br /&gt;
* Code submissions by Niemah&lt;br /&gt;
[[File:Presidents Shuffle bar plot.png|thumbnail]]&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&#039;bash&#039;&amp;gt;&lt;br /&gt;
library(tidyverse)&lt;br /&gt;
p &amp;lt;- read_tsv(&amp;quot;Presidents.txt&amp;quot;, col_names = FALSE )&lt;br /&gt;
colnames(p) &amp;lt;- c(&amp;quot;order&amp;quot;, &amp;quot;name&amp;quot;, &amp;quot;year.elected&amp;quot;)&lt;br /&gt;
true.order &amp;lt;- p$name&lt;br /&gt;
output &amp;lt;- vector(&amp;quot;double&amp;quot;, length = 10000)&lt;br /&gt;
for(i in 1:10000) {&lt;br /&gt;
  random.order &amp;lt;- sample(p$name)&lt;br /&gt;
  output[[i]] &amp;lt;- sum(true.order == random.order)&lt;br /&gt;
}&lt;br /&gt;
x &amp;lt;- table(output)&lt;br /&gt;
df &amp;lt;- tibble(run = 1:1e4, matches = output)&lt;br /&gt;
df %&amp;gt;% ggplot(aes(x=matches)) + geom_bar(fill = &amp;quot;lightgreen&amp;quot;) + theme_minimal()&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&#039;python&#039;&amp;gt;&lt;br /&gt;
import pandas as pd&lt;br /&gt;
import numpy as np&lt;br /&gt;
import matplotlib.pyplot as plt&lt;br /&gt;
import random&lt;br /&gt;
&lt;br /&gt;
# read file&lt;br /&gt;
file = r&#039;/content/Presidents.txt&#039;&lt;br /&gt;
dp = pd.read_table(&#039;Presidents.txt&#039;, header=None, names=[&#039;order&#039;, &#039;name&#039;, &#039;year.elected&#039;])&lt;br /&gt;
print(dp)&lt;br /&gt;
names_list = list(dp.name)&lt;br /&gt;
print(names_list)&lt;br /&gt;
&lt;br /&gt;
# simulate by 10000 permutations&lt;br /&gt;
output = [] # initialize an empty vector to store the match numbers&lt;br /&gt;
for x in range(10000):&lt;br /&gt;
  names_permuted = np.random.permutation(dp.name)&lt;br /&gt;
  num_match = sum(names_list == names_permuted)&lt;br /&gt;
  output.append(num_match)&lt;br /&gt;
print(output)&lt;br /&gt;
&lt;br /&gt;
# tabulate counts&lt;br /&gt;
table_match = {} # initialize a dict to store counts&lt;br /&gt;
for num in output:&lt;br /&gt;
  if num in table_match:&lt;br /&gt;
    table_match[num] += 1&lt;br /&gt;
  else: # initialize if the number is first seen&lt;br /&gt;
    table_match[num] = 0&lt;br /&gt;
print(table_match)&lt;br /&gt;
&lt;br /&gt;
# plot&lt;br /&gt;
fig = plt.figure()&lt;br /&gt;
ax = fig.add_axes([0,0,1,1])&lt;br /&gt;
x = []&lt;br /&gt;
y = []&lt;br /&gt;
for match, cts in table_match.items():&lt;br /&gt;
    x.append(match)&lt;br /&gt;
    y.append(cts)&lt;br /&gt;
ax.bar(x,y)&lt;br /&gt;
ax.set_title(&amp;quot;Number of President Matches&amp;quot;)&lt;br /&gt;
ax.set_xlabel(&#039;Matches&#039;)&lt;br /&gt;
ax.set_ylabel(&#039;Count&#039;)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Week 4==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
!  Entropy plots (Winston &amp;amp; Anh) !! Fitness plots &lt;br /&gt;
|-&lt;br /&gt;
| [[File:Entropy.png|thumbnail]] || [[File:fit.png|thumbnail]] &lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Week 6==&lt;br /&gt;
* Slides by Anh: [[File:Structure and Mutations of SARS-CoV-2 Spike Protein Anh Pham.pdf|thumbnail]]&lt;br /&gt;
* Slides by Winston: [[File:Binding of Human ACE2 and RBD of Omicron.pdf| thumbnail | Presentation by Winston]]&lt;br /&gt;
==Week 7==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Fitness landscapes of binary strings&lt;br /&gt;
|-&lt;br /&gt;
| &lt;br /&gt;
[[File:Landscape.png|thumbnail| three distributions of fitness of mutations (by Anh &amp;amp; Winston)]]&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Reading list==&lt;br /&gt;
* [https://pages.cs.wisc.edu/~dyer/cs540/handouts/info-theory-primer.pdf A primer for information theory]&lt;br /&gt;
* A somewhat advanced introduction of information theory by [https://sites.santafe.edu/~simon/it.pdf Simon DeDeo (2018)]&lt;br /&gt;
* Data ≠ Information: Scharf (2021). The Ascent of Information: Books, Bits, Genes, Machines, and Life&#039;s Unending Algorithms. [https://www.amazon.com/Ascent-Information-Machines-Unending-Algorithm/dp/0593087240 Amazon link]&lt;br /&gt;
* The &amp;quot;It&#039;s the song not the singer&amp;quot; (ITSNTS) theory of selection unit: [https://pubmed.ncbi.nlm.nih.gov/29581311/ Doolittle &amp;amp; Inkpen (2018).] &amp;quot;Processes and patterns of interaction as units of selection: An introduction to ITSNTS thinking&amp;quot;, PNAS.&lt;br /&gt;
* Evolution is an AI machine: &lt;br /&gt;
** [https://www.oreilly.com/radar/open-endedness-the-last-grand-challenge-youve-never-heard-of/ Stanley, Lehman, and Soros (2017).] &amp;quot;Open-endedness: The last grand challenge you’ve never heard of - While open-endedness could be a force for discovering intelligence, it could also be a component of AI itself.&amp;quot; O&#039;Reily&lt;br /&gt;
** Novelty search algorithm: towards evolutionary complexity. a complexification algorithm that can be integrated into any evolutionary computation (EC)&lt;br /&gt;
** [https://www.oreilly.com/radar/open-endedness-a-new-grand-challenge-for-ai/ O&#039;Reily talk video clip by Ken Stanley (2019)]&lt;br /&gt;
** [http://eplex.cs.ucf.edu/ EPLEX group]&lt;br /&gt;
** [https://nbenko1.github.io/#/ Evolutionary art: picbreeder]&lt;br /&gt;
** [http://eplex.cs.ucf.edu/noveltysearch/userspage/#howtoimplement Implementation: add a vector to characterize behavior (not fitness)]&lt;br /&gt;
** NEAT paper: [https://www.cs.swarthmore.edu/~meeden/cs81/f17/projects/JeffKyle.pdf Novak and Richmond-Crosset (2017)]&lt;br /&gt;
** Evolutionary creativity: [[https://direct.mit.edu/artl/article/26/2/274/93255/The-Surprising-Creativity-of-Digital-Evolution-A|Lehman, Clune, and Misevic (2020). Surprising creativity of digital evolution. Alife]]&lt;br /&gt;
* An algorithmic definition of individuality: [https://pubmed.ncbi.nlm.nih.gov/32212028/ Krakauer et al (2020)]. &amp;quot;The information theory of individuality&amp;quot;. Theory in BioSci&lt;br /&gt;
* Evolution towards complexification: [https://aip.scitation.org/doi/abs/10.1063/1.3643064 Krakauer (2011)]. &amp;quot;Darwinian demons, evolutionary complexity, and information maximization&amp;quot;. Chaos.&lt;br /&gt;
* &amp;quot;Aliveness&amp;quot;: [https://link.springer.com/article/10.1007%2Fs12064-020-00331-5 Kim  et al (2021)]. &amp;quot;Informational architecture across non-living and living collectives&amp;quot;. Theory in BioSci&lt;br /&gt;
* [https://www.pnas.org/content/97/9/4463.long Adami et al (2000). Evolution of biological complexity. PNAS]&lt;br /&gt;
* Evolution-based protein designs &amp;amp; structure-function analysis&lt;br /&gt;
** To increase stability: [https://doi.org/10.1016/j.molcel.2016.06.012 Goldenzweig et al (2021).]&lt;br /&gt;
** Full-spectrum mutation scan: [https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3638900/ Whitehead et al (2012).]&lt;br /&gt;
** Full-spectrum mutation scan of 4 sites GB1 protein, adaptive fitness landscape with higher-order epistasis: [https://elifesciences.org/articles/16965 Wu et al (2016) eLife]&lt;br /&gt;
** Full mutation scan of influenza antigenic drift sites: [https://www.nature.com/articles/s41467-020-15102-5 Wu et al (2020)]; [https://pubmed.ncbi.nlm.nih.gov/34878407/ Wang et al (2021)]&lt;br /&gt;
** A review of SARS-CoV-2 S protein structure: effects of common mutations. [https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8673470/ Mehara &amp;amp; Kepp (2021)] [[File:Structure and Mutations of SARS-CoV-2 Spike Protein Anh Pham.pdf|thumbnail]]&lt;br /&gt;
** Computational modeling of Omicron mutations: [https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8811907/ Kim et al (2022)] [[File:Binding of Human ACE2 and RBD of Omicron.pdf| thumbnail | Presentation by Winston]]&lt;br /&gt;
* Predicting fitness of haplotypes from individual mutations; epistasis&lt;br /&gt;
** [https://journals.plos.org/plospathogens/article?id=10.1371/journal.ppat.1009669 Pedruzzi &amp;amp; Rouzine (2021). PLoS Pathogen.] &amp;quot;An evolution-based high-fidelity method of epistasis measurement: theory and application to influenza&amp;quot;: a method to distinguish covariance by linkage and covariance by epistasis (selection), with &amp;quot;tri-way haplotype frequency&amp;quot;&lt;br /&gt;
** Universal exponential Distribution of Fitness Effect (DFE) of mutations, resistant to linkage and epistasis: [https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1008822 Barlukova &amp;amp; Rouzine (2021). PLoS Comp Biol.]&lt;br /&gt;
** Estimate fitness landscape of gp160, a large HIV variable surface protein: [https://www.pnas.org/doi/full/10.1073/pnas.1717765115 Louie et al (2018).] PNAS. &amp;quot;Fitness landscape of the human immunodeficiency virus envelope protein that is targeted by antibodies&amp;quot;&lt;br /&gt;
** Fitness landscape of Ribozyme: [https://journals.plos.org/plosbiology/article?id=10.1371/journal.pbio.3000300 Bendixsen et al (2019).] PLoS Biology. &amp;quot;Genotype network intersections promote evolutionary innovation&amp;quot;&lt;br /&gt;
** Fitness landscape in Dengue virus, based on ultra-RNA sequencing &amp;amp; growth measurements, indicating numerous deleterious mutations. [https://elifesciences.org/articles/61921 Dolan et al.(2021) eLife]&lt;br /&gt;
** Fitness landscape of SARS-CoV-2: genetic and non-genetic components. [https://academic.oup.com/ve/article/7/2/veab073/6363035 Kepler et al (2021). Virus Evolution.]&lt;br /&gt;
** Predicting mutation fitness of SARS-Cov-2: Obermeyer et al (2022). Analysis of 6.4 million SARS-CoV-2 genomes identifies mutations associated with fitness. [https://www.science.org/doi/10.1126/science.abm1208 Science] &lt;br /&gt;
** Deep mutation scan of SARS-CoV-2 RBD-binding domain: [https://www.science.org/doi/epdf/10.1126/science.abo7896 Starr et al (2022). Science]&lt;br /&gt;
** Song &amp;amp; Zhang (2022). &amp;quot;Unbiased inference of the fitness landscape ruggedness from imprecise fitness estimates&amp;quot;. 3 fitness landscapes: NK, RMF, and polynomial fitness models. [https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9018209/ PubMed link]. [https://github.com/song88180/fitness-landscape-error Github repository]&lt;br /&gt;
&lt;br /&gt;
==Tookits==&lt;br /&gt;
* [https://sites.santafe.edu/~simon/page7/page7.html THOTH: a Python package for Shanon information]&lt;br /&gt;
* [https://github.com/weigangq/cov-db CoV simulator on Github]&lt;br /&gt;
* [https://www.laurentluce.com/posts/solving-mazes-using-python-simple-recursivity-and-a-search/ A maze solver using novelty search (in Python)]&lt;br /&gt;
* [https://github.com/yaricom/goNEAT_NS Github for NEAT-based maze solver]: behavioral vector: ending position:&lt;br /&gt;
** &amp;quot;Therefore, because what is important in a maze is where the solving agent ends, for the maze domain, the behavior of a navigator is defined as its ending position. The novelty metric is then the N-nearest neighbor distance novelty between the ending positions of all known solving agents.&amp;quot;&lt;br /&gt;
* A collection of divergence and diversity search: [https://github.com/DanieleGravina/divergence-and-quality-diversity github]: novelty search, surprise search, curiosity search, quality diversity, etc&lt;br /&gt;
* [https://github.com/CodeReclaimers/neat-python NEAT-Python] &amp;amp; various forks on Github&lt;br /&gt;
* Novelty search by [https://github.com/peter-ch/MultiNEAT multiNEAT], python implementation&lt;br /&gt;
* Visualization of fitness landscape &amp;amp; path with R&lt;br /&gt;
** [https://r-charts.com/correlation/contour-plot/ Basic contour function] &lt;br /&gt;
** [https://plotly.com/r/contour-plots/ Contour map by R plotly]&lt;br /&gt;
* Links for Bayesian nested models&lt;br /&gt;
** with rstan: [https://biologyforfun.wordpress.com/2016/12/08/crossed-and-nested-hierarchical-models-with-stan-and-r/ on WorldPress]&lt;br /&gt;
** [https://www.r-bloggers.com/2016/12/crossed-and-nested-hierarchical-models-with-stan-and-r/ | another example on r-bloggers]&lt;br /&gt;
** [https://www.r-bloggers.com/2016/12/crossed-and-nested-hierarchical-models-with-stan-and-r/ Nested model with categorical predictors]&lt;br /&gt;
&lt;br /&gt;
=Season IV. Classification using Machine Learning (Summer 2021)=&lt;br /&gt;
* Textbook: Aurélien Géron (2019). Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems, 2nd Edition. [https://www.amazon.com/gp/product/1492032646/ref=ppx_yo_dt_b_asin_title_o00_s00?ie=UTF8&amp;amp;psc=1 Amazon link]&lt;br /&gt;
* Set up Python work environment with [https://docs.conda.io/en/latest/miniconda.html mini-conda]&lt;br /&gt;
* We will be using Jupyter-notebook (part of mini-codon installation) to share codes&lt;br /&gt;
==Week 1. MNIST dataset (Chapter 3)==&lt;br /&gt;
# Dataset loading and display (pg 85-87): Jackie, Niemah, and Hannah&lt;br /&gt;
# Binary classifier&lt;br /&gt;
## Cross-validation (pg 89-90): Roman, etc&lt;br /&gt;
## Confusion matrix (pg 90-92): &lt;br /&gt;
## Precision and recall (pg 92-97)&lt;br /&gt;
## ROC curve (pg. 97-100)&lt;br /&gt;
# Multiclass classification (pg 100-108): Brian, etc&lt;br /&gt;
&lt;br /&gt;
==Week 2. K-means clustering (Chapter 9)==&lt;br /&gt;
# 2D simulated dataset&lt;br /&gt;
# Image recognition&lt;br /&gt;
# MNIST dataset&lt;br /&gt;
==Week 3. Exercises (pg.275-276, Chapter 9)==&lt;br /&gt;
# Exercise 10: Facial recognition (Olivetti faces dataset, with k-means)&lt;br /&gt;
# Exercise 11. Facial recognition (semi-supervised learning&lt;br /&gt;
&lt;br /&gt;
=Notes on Origin of Life (Spring 2020)=&lt;br /&gt;
# [https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0224552 Attie et al (2019). Genetic Code optimized as a traveling salesman problem]&lt;br /&gt;
# [https://itsatcuny.org/calendar/self-organizing-systems-and-the-origin-of-life GC Origin of Life Seminar (2/19/2021)]&lt;br /&gt;
# [https://www.cell.com/current-biology/fulltext/S0960-9822(15)00681-8?_returnURL=https%3A%2F%2Flinkinghub.elsevier.com%2Fretrieve%2Fpii%2FS0960982215006818%3Fshowall%3Dtrue Pressman et al (2015). Review: RNA World (ribozyme as origin of genetic code)]&lt;br /&gt;
# A perspective: [https://www.nature.com/articles/s42256-020-00278-8 Miikkulainen, R., Forrest, S. A biological perspective on evolutionary computation. Nat Mach Intell 3, 9–15 (2021)]&lt;br /&gt;
&lt;br /&gt;
=Season III. Summer 2018 (Theme: Evolutionary Computing)=&lt;br /&gt;
==Week 1. Introduction &amp;amp; Motivating Examples==&lt;br /&gt;
# Genetic Arts&lt;br /&gt;
## [http://picbreeder.org Go to picbreeker website] &amp;amp; evolve using &amp;quot;branch&amp;quot; method&lt;br /&gt;
## Evolve 3D art: [http://endlessforms.com Endless Forms]&lt;br /&gt;
## CPPN-NEAT Algorithm: Compositional Pattern Producing Networks (CPPNs)-NeuroEvolution of Augmenting Topologies (NEAT); [https://www.ncbi.nlm.nih.gov/pubmed/20964537 PicBreeder paper]; or [http://campbellssite.com/papers/secretan_chi08.pdf a PDF version]&lt;br /&gt;
# [https://www.nature.com/articles/s41586-018-0102-6 NeuroEvolution (by DeepMind team)]&lt;br /&gt;
# Robotic snake (and other soft robots)&lt;br /&gt;
## A controller built with physics laws difficult (too many parameters)&lt;br /&gt;
## Simulation with evolutionary computing: &lt;br /&gt;
### (Genotype) A list of 13 commands (one for each segment; each being a neural net, with 25 inputs and one output of joint angles)&lt;br /&gt;
### (Phenotype) Fitness function: total displacement&lt;br /&gt;
### Algorithm: [https://medium.com/@devonfulcher3/the-map-elites-algorithm-finding-optimality-through-diversity-def6dcbc0f5b MAP-Elites]; [https://www.nature.com/articles/nature14422 Nature paper]&lt;br /&gt;
### Approach: training with simulated data&lt;br /&gt;
# Robotic Knightfish&lt;br /&gt;
## [https://www.youtube.com/watch?v=3XjgZbs0t2g Youtube demo]&lt;br /&gt;
## Algorithm: [https://en.wikipedia.org/wiki/CMA-ES CMA-ES (Covariance Matrix Adaptation Evolution Strategy)]&lt;br /&gt;
### Genotype: 15 variables (sinusoidal wave function or Fourier Series)&lt;br /&gt;
### Phenotype/Fitness: speed&lt;br /&gt;
## Implementation: DEAP&lt;br /&gt;
# Compositional Protein design (CPD)&lt;br /&gt;
## Genotype: side-chain configuration determined by [https://www.ncbi.nlm.nih.gov/pubmed/8464064 rotamer library]&lt;br /&gt;
## Phenotype/Fitness: Rosetta energy function &amp;amp; functional (e.g., binding affinity)&lt;br /&gt;
## Fitness landscape: each node is a protein structure, each edge represent connections/relatedness&lt;br /&gt;
# Simulation-based optimization: [http://simopt.org/ Problem Sets]&lt;br /&gt;
# Self-evolving software([http://geneticprogramming.com/ Genetic Programming])&lt;br /&gt;
&lt;br /&gt;
==Week 2. Toy Problem: OneMax Optimization==&lt;br /&gt;
# Problem: Create a list of L random bits (0 or 1). Evolve the list until the fitness reaches the maximum value (i.e., contains only 1&#039;s)&lt;br /&gt;
# Neutral Evolution:&lt;br /&gt;
## Create a vector of L random bits (e.g, L=20). Hint for creating a random vector of 0&#039;s and 1&#039;s in R: &amp;lt;code&amp;gt;ind&amp;lt;-sample(c(0,1), prob=c(0.5.0.5), replace=T, size=20)&amp;lt;/code&amp;gt;. Biologically, this vector represents a single haploid genome with 20 loci, each with two possible alleles (0 or 1).&lt;br /&gt;
## Create a population of N=100 such individuals. Hint: creating a list of vectors in R: &amp;lt;code&amp;gt;pop &amp;lt;- lapply(1:100, function(x){&amp;lt;insert sample() function above&amp;gt;})&amp;lt;/code&amp;gt;&lt;br /&gt;
## For each generation, each individual reproduces 10 gametes with mutation (with the probability for bit flip: mu = 1/L = 0.1). Hint: write a mutation function that takes an individual as input and outputs a mutated gamete. Use the &amp;quot;broken stick&amp;quot; algorithm to implement mutation rate: &amp;lt;code&amp;gt;cutoff &amp;lt;- runif(1); ifelse(cutoff &amp;lt;= mu, flip-bits, no-flip)&amp;lt;/code&amp;gt;&lt;br /&gt;
## Take a random sample of N=100 gametes into the next generation &amp;amp; repeat above&lt;br /&gt;
## Plot mean, minimum, and maximum fitness over generations&lt;br /&gt;
## Plot diversity statistics over generation (including average allelic heterozygosity per locus as well as haplotype heterozygosity). Hint: write two functions for these heterozygosity&lt;br /&gt;
# Add natural selection (proportional scheme)&lt;br /&gt;
## Individuals reproduce with fitness proportional to the total number of 1&#039;s.&lt;br /&gt;
## Iterate until a population contains at least one individual with all 1&#039;s&lt;br /&gt;
## Plot mean, minimum, and maximum fitness over generations&lt;br /&gt;
## Plot diversity (expectation: decreasing)&lt;br /&gt;
# Add natural selection (tournament scheme)&lt;br /&gt;
## Randomly selecting n=5 individuals and allow the fittest one to make gametes&lt;br /&gt;
## Plot mean, minimum, and maximum fitness over generations (Expectation: faster optimization)&lt;br /&gt;
## Plot diversity (expectation: decreasing fasters)&lt;br /&gt;
# Add crossover&lt;br /&gt;
## Hint: write a crossover function&lt;br /&gt;
## Does it reach optimization faster?&lt;br /&gt;
# Code submissions&lt;br /&gt;
## [https://github.com/JohnDi0505/MonteCarlo_Simulation-Biostats/blob/master/OneMax%20Optimization.ipynb Python Notebook by John]&lt;br /&gt;
## [http://rpubs.com/weigang/407979 R code by Weigang]&lt;br /&gt;
## [http://rpubs.com/ChrisNSP/408534 R code by Panlasigui]&lt;br /&gt;
## [https://github.com/Inhenn/Toy_Problem_Bio/blob/master/Bio_Toy.ipynb Python Notebook by Yinheng]&lt;br /&gt;
## [http://rpubs.com/desiree/409200 RPub by Desiree]&lt;br /&gt;
==Week 3. Python &amp;amp; R Packages for evolutionary computation==&lt;br /&gt;
* [http://deap.readthedocs.io/en/master/ DEAP: a Python package for Evolutionary Computing]&lt;br /&gt;
** Look under [http://deap.readthedocs.io/en/master/examples/index.html Examples] to repeat the OneMax code&lt;br /&gt;
* [https://cran.r-project.org/web/packages/GA/vignettes/GA.html GA: An R package for genetic programming]&lt;br /&gt;
** Reference 1 (Examples in Section 4). [https://www.jstatsoft.org/v53/i04/ Scrucca, L. (2013) GA: A Package for Genetic Algorithms in R.]&lt;br /&gt;
** Reference 2 (for Advanced applications). [https://journal.r-project.org/archive/2017/RJ-2017-008 Scrucca, L. (2017) On some extensions to GA package: hybrid optimisation, parallelisation and islands evolution. ]&lt;br /&gt;
** OneMax code &amp;amp; plots (Weigang)&lt;br /&gt;
** Example 4.1a. One variable optimization: f(x) = |x| + cos(x) (Muhammad and Desiree)&lt;br /&gt;
** Example 4.1b. One variable optimization: f(x) = (x2 + x) cos(x)&lt;br /&gt;
** Example 4.2.  Two-parameter optimization: f(x;y) = 20 + x^2 + y^2 * 10(cos(2x) + cos(2y)) (Muhammad and Desiree)&lt;br /&gt;
** Example 4.3.  Curve-fitting: tree growth&lt;br /&gt;
** Example 4.7.  Constrained optimization: Knapsack Problem&lt;br /&gt;
** Example 4.8.  Combinatorial optimization: Traveling Salesman Problem (Brian)&lt;br /&gt;
** Advanced application 1. Stock portfolio (Hybrid algorithms)&lt;br /&gt;
** Advanced application 2. Parallelization&lt;br /&gt;
** Advanced application 3. Island model&lt;br /&gt;
* Code submissions&lt;br /&gt;
** [http://rpubs.com/weigang/410426 rPub for OneMax (by Weigang)]&lt;br /&gt;
** One &amp;amp; Two-dimensional functional optimization with GA (by Desiree &amp;amp; Mohamud) : [http://rpubs.com/desireepante/411065 Entropy function]&lt;br /&gt;
** [https://github.com/Inhenn/Knapsack-Problem-Using-DEAP/blob/master/DEAP3.ipynb OneMax and Knapsack Problem with DEAP (by Yinheng)]&lt;br /&gt;
** Traveling Salesman Problem with GA (by Brian)&lt;br /&gt;
&lt;br /&gt;
==Week 4. Multiplex Problem==&lt;br /&gt;
==Week 5. Genetic Programming==&lt;br /&gt;
&lt;br /&gt;
=Season II. Summer 2017 (Theme: Machine Learning)=&lt;br /&gt;
==Week 1. Introduction &amp;amp; the backprop algorithm==&lt;br /&gt;
[[File:Iris-box3.png|thumbnail]] [[File:Iris-box4.png|thumbnail]]&lt;br /&gt;
# Problem: Classification/Clustering/Predication of flower species (a total of 3 possible species in the sample data set) based on four phenotypic traits/measurements&lt;br /&gt;
# The &amp;quot;iris&amp;quot; data set: exploratory data analysis with visualization &amp;amp; descriptive statistics: &amp;lt;code&amp;gt;data(&amp;quot;iris&amp;quot;); plot()&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;summary()&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;qqnorm(); qqline(); hist()&amp;lt;/code&amp;gt;; normalization with &amp;lt;code&amp;gt;scale()&amp;lt;/code&amp;gt; (Roy)&lt;br /&gt;
# Mathematics of backpropagating errors to neural connections (Oliver)&lt;br /&gt;
## Objective/optimization function measuring difference between target (&amp;lt;code&amp;gt;t&amp;lt;/code&amp;gt;, expected) and neural activity &amp;lt;code&amp;gt;y&amp;lt;/code&amp;gt;: &amp;lt;code&amp;gt;G=Sum{t*log(y)+(1-t)log(1-y)}&amp;lt;/code&amp;gt; (this is known as the &amp;quot;cross-entropy&amp;quot; error function; the other alternative is &amp;quot;minimal squared error (MSE)&amp;quot;), which has the gradient in the simple form of &amp;lt;code&amp;gt;g=-(t-y)x, where x is the input&amp;lt;/code&amp;gt;. The objective function is minimized when weights are updated by the gradient. Error-minimization by MSE has similar effects but harder to calculate.&lt;br /&gt;
## Learning algorithm is presented&lt;br /&gt;
&lt;br /&gt;
==Week 2. Traditional approaches to multivariate clustering/classification==&lt;br /&gt;
# Dimension reduction with Multidimensional Scaling &amp;lt;code&amp;gt;cmdscale()&amp;lt;/code&amp;gt; [http://rpubs.com/meibyderp/281616 Mei&#039;s rNoteBook]. &lt;br /&gt;
# Dimension reduction with Principal Component Analysis&amp;lt;code&amp;gt;princomp()&amp;lt;/code&amp;gt;. [http://rpubs.com/ssipa/281715 Sipa&#039;s rNoteBook]&lt;br /&gt;
# Multivariate clustering with Hierarchical Clustering &amp;lt;code&amp;gt;hclust()&amp;lt;/code&amp;gt; (Saymon)&lt;br /&gt;
# Multivariate clustering with k-means &amp;lt;code&amp;gt;kmeans()&amp;lt;/code&amp;gt; [http://rpubs.com/roynunez/281963 Roy&#039;s rNoteBook]&lt;br /&gt;
# Classification based on logistic regression &amp;lt;code&amp;gt;glm()&amp;lt;/code&amp;gt;, linear discriminatory analysis &amp;lt;code&amp;gt;lda()&amp;lt;/code&amp;gt;, &amp;amp; k-nearest neighbor &amp;lt;code&amp;gt;library(class); knn()&amp;lt;/code&amp;gt; (John)&lt;br /&gt;
# Modern, non-linear classifiers: Decision Trees (DT), Support Vector Machines (SVM), and Artificial Neural Networks (ANN) (Brian)&lt;br /&gt;
&lt;br /&gt;
==Week 3. Single-neuron classifier==&lt;br /&gt;
# Algorithm: [[File:Ml-image-1a.jpg|thumbnail]]&lt;br /&gt;
## Read input data with &amp;lt;code&amp;gt;N=150 flowers&amp;lt;/code&amp;gt;. Reduce to a matrix &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; with two traits (use trait 1 &amp;amp; 3) and two species (use rows 51-150, the last two species, skip the first species [easy to separate]) for simplicity. Create a target vector &amp;lt;code&amp;gt;t &amp;lt;- c(rep(0,50),rep(1,50))&amp;lt;/code&amp;gt; indicating two species&lt;br /&gt;
## Initialize the neuron with two random weights &amp;lt;code&amp;gt;w&amp;lt;-runif(2, 1e-3, 1e-2)&amp;lt;/code&amp;gt; and one random bias &amp;lt;code&amp;gt;b&amp;lt;-runif(1)&amp;lt;/code&amp;gt;&lt;br /&gt;
## Neuron activation with &amp;lt;code&amp;gt;k=2&amp;lt;/code&amp;gt; connection weights: &amp;lt;code&amp;gt;a=sum(x[k] * w[k])&amp;lt;/code&amp;gt;&lt;br /&gt;
## Neuron activity/output: &amp;lt;code&amp;gt;y=1/(1+exp(-a-b))&amp;lt;/code&amp;gt; (This logistic function ensures output values are between zero and one)&lt;br /&gt;
## Learning rules: learning rate &amp;lt;code&amp;gt;eta=0.1&amp;lt;/code&amp;gt;, backpropagate error (&amp;quot;e&amp;quot;) to get two updated weights (for individual &amp;lt;code&amp;gt;i&amp;lt;/code&amp;gt;, feature &amp;lt;code&amp;gt;k&amp;lt;/code&amp;gt;): &amp;lt;code&amp;gt;e[i]=t[i]-y[i]; g[k,i]= -e[i] * x[k,i]; g.bias[i] = -e[i] for bias&amp;lt;/code&amp;gt;; Batch update weights &amp;amp; bias: &amp;lt;code&amp;gt;w[k]=w[k] - eta * sum(g[k,i]); b = b - eta * sum(g.bias[i])&amp;lt;/code&amp;gt; (same rule for the bias parameter &amp;lt;code&amp;gt;b&amp;lt;/code&amp;gt;); Repeat/update for &amp;lt;code&amp;gt;L=1000 epochs (or generations)&amp;lt;/code&amp;gt;&lt;br /&gt;
## Output: weights and errors for each epoch&lt;br /&gt;
# Evaluation:&lt;br /&gt;
## Plot changes of weights over epoch&lt;br /&gt;
## Use the last weights and bias to predict species&lt;br /&gt;
## Compare prediction with target to get accuracy&lt;br /&gt;
## Plot scatter plot (x2 vs x1) and add a line using the weights &amp;amp; bias at epoch=1,50,100,200,500, 1000: &amp;lt;code&amp;gt;a=w1*x1 + w2*x2 + b; a=0&amp;lt;/code&amp;gt;. The lines should show increasing separation of the two species.&lt;br /&gt;
# Code submissions&lt;br /&gt;
## Roy: [http://rpubs.com/roynunez/283699 rPubs notebook]&lt;br /&gt;
## John: [https://github.com/JohnDi0505/MonteCarlo_Simulation-Biostats/blob/master/Neuron_Networks/Single_Neuron_Classification.py Python code on github]&lt;br /&gt;
## Mei: [http://rpubs.com/meibyderp/284498 rPubs notebook]&lt;br /&gt;
## Brian: [http://rpubs.com/drtwisto/286938 rPubs notebook]&lt;br /&gt;
## Weigang: [http://rpubs.com/weigang/279898 rPubs notebook]&lt;br /&gt;
# Questions for future exploration&lt;br /&gt;
## How to avoid over-fitting by regularization (a way to penalize model complexity by adding a weight decay parameter &amp;lt;code&amp;gt;alpha&amp;lt;/code&amp;gt;)&lt;br /&gt;
## Bayesian confidence interval of predictions (with Monte Carlo simulation)&lt;br /&gt;
## Limitations: equivalent to PCA (linear combination of features); adding a hidden layer generalize the neural net to be a non-linear classifier&lt;br /&gt;
&lt;br /&gt;
==Week 4. Single-layer, multiple-neuron learner==&lt;br /&gt;
[[File:Multiple-neuron-learner.png|thumbnail]]&lt;br /&gt;
# Predict all three species (with &amp;quot;one-hot&amp;quot; coding) using all four features: e.g., 100 for species 1, 010 for species 2, and 001 for species 3. Create 3 neurons, each one outputting one digit.&lt;br /&gt;
# Use three neurons, each accepts 4 inputs and output 1 activity&lt;br /&gt;
# Use softmax to normalize the final three output activities&lt;br /&gt;
# Code submissions&lt;br /&gt;
## John: &lt;br /&gt;
### [http://www.kdnuggets.com/2016/07/softmax-regression-related-logistic-regression.html reference this webpage]&lt;br /&gt;
### [https://github.com/JohnDi0505/MonteCarlo_Simulation-Biostats/blob/master/Neuron_Networks/Multiple_Neuron_Classification.py Python code]&lt;br /&gt;
### TensorFlow code&lt;br /&gt;
## Roy: [http://rpubs.com/roynunez/288023 rpubs notebook]&lt;br /&gt;
## Mei: [http://rpubs.com/meibyderp/288400 rPubs Notebook]&lt;br /&gt;
## Weigang: [http://rpubs.com/weigang/287025 rPubs notebook]&lt;br /&gt;
# Challenge: Implementation with TensorFlow, an Open Source python machine learning library by Google Deep Mind team. Follow [https://www.tensorflow.org/get_started/mnist/beginners this example of softmax regression]&lt;br /&gt;
# A biological application for the summer: identify SNPs associated with biofilm/swarming behavior in Pseudomonas&lt;br /&gt;
&lt;br /&gt;
==Week 5. Multi-layer (&amp;quot;Deep&amp;quot;) Neuron Network==&lt;br /&gt;
[[File:Iris-nnet.png|thumbnail]]&lt;br /&gt;
[[File:Nnet-overfitting.png|thumbnail|An investigation of under-fitting (at N=1) &amp;amp; over-fitting (at N=3 nodes). N=0 implies linear fitting.]]&lt;br /&gt;
# Input layer: 4 nodes (one for each feature)&lt;br /&gt;
# Output layer: 3 nodes (one for each species)&lt;br /&gt;
# Hidden layer: 2 hidden nodes&lt;br /&gt;
# Advantage: allows non-linear classification; using hidden layers (&amp;quot;convoluted neural net&amp;quot;) is able to capture high-order patterns&lt;br /&gt;
# Implementation: I&#039;m not going to hard-code the algorithm from scratch. The figure was produced using the R package with the following code: &amp;lt;code&amp;gt;library(nnet); library(neuralnet); targets.nn &amp;lt;- class.ind(c(rep(&amp;quot;setosa&amp;quot;,50), rep(&amp;quot;versicolor&amp;quot;,50), rep(&amp;quot;virginica&amp;quot;,50))) # 1-of-N encoding; iris.net &amp;lt;- neuralnet(formula = setosa + versicolor + virginica ~ Sepal.Length + Sepal.Width + Petal.Length + Petal.Width, data = training.iris, hidden = 2, threshold = 0.01, linear.output = T); plot(iris.net, rep=&amp;quot;best&amp;quot;)&amp;lt;/code&amp;gt;. It achieved 98.0% accuracy.&lt;br /&gt;
# Deep neural net allows non-linear, better fitting, but we don&#039;t want over-fitting by adding more hidden layers (or more neurons in the hidden layer). Identify under- and over-fitting with the following procedure:&lt;br /&gt;
## Randomly sample 100 as training set and the remaining as target. Repeat 100 times&lt;br /&gt;
## For each sample, plot accuracy for the training set, as well as accuracy for the target&lt;br /&gt;
## Find the point with the right balance of under- and over-fitting&lt;br /&gt;
&lt;br /&gt;
==Week 6. Unsupervised Neural Learner: Hopfield Networks==&lt;br /&gt;
# Hebbian model of memory formation (MacKay, Chapter 42). Preparatory work:&lt;br /&gt;
## Construct four memories, each for a letter (&amp;quot;D&amp;quot;, &amp;quot;J&amp;quot;, &amp;quot;C&amp;quot;, and &amp;quot;M&amp;quot;) using a 5-by-5 grid, with &amp;quot;-1&amp;quot; indicating blank space, and &amp;quot;1&amp;quot; indicating a pixel. Flatten the 5-by-5 matrix to a one-dimensional 1-by-25 vector (tensor). &lt;br /&gt;
## Write two functions, one to show a letter &amp;lt;code&amp;gt;show.letter(letter.vector)&amp;lt;/code&amp;gt;, which draws a pixel art of letters (print a blank if if -1, a &amp;quot;x&amp;quot; if 1), another to mutate the letter &amp;lt;code&amp;gt;mutate(letter.vector, number.pixel.flips)&amp;lt;/code&amp;gt;&lt;br /&gt;
# [https://en.wikipedia.org/wiki/Hopfield_network Hopfield Network]&lt;br /&gt;
## Store the four memories into a weight matrix, which consists of symmetric weights between neurons i and neuron j, i.e., w[i,j] = w[j,i]. We will use a total of 25 neurons, one for each pixel.&lt;br /&gt;
## First, combine the four vectors (one for each letter) into a matrix: &amp;lt;code&amp;gt;x &amp;lt;- matrix(c(letter.d, letter.j, letter.c, letter.m), nrow = 4, byrow = T)&amp;lt;/code&amp;gt;&lt;br /&gt;
## Second, calculate weight matrix by obtaining the outer product of a matrix multiplication: &amp;lt;code&amp;gt;w &amp;lt;- t(x) %*% x&amp;lt;/code&amp;gt;. Set diagonal values to be all 0&#039;s (i.e., to remove all self-connections): &amp;lt;code&amp;gt;for (i in 1:25) { w[i,i] = 0 }&amp;lt;/code&amp;gt;. This implements [https://en.wikipedia.org/wiki/Hebbian_theory Hebb learning], which translates correlations into strength of connections quantified by weights: large positive weights indicate mutual stimulation (e.g., 1 * 1 = 1 [to wire/strengthen the connection of co-firing neurons, and ...], -1 * -1 = 1 [to wire/strengthen the connection of co-inhibitory neurons as well]), large negative weights indicate mutual inhibition (e.g., 1 * -1 = -1 [to unwire/disconnect oppositely activated neurons]), and small weights indicate a weak connection (e.g., 0 * 1 = 0 [to weaken connections between neurons with uncorrelated activities, but do not unwire them]). (0, 1, and -1 being values of neuron activities)&lt;br /&gt;
## Third, implement the learning rule by updating activity for each neuron, sequentially (asynchronously): &amp;lt;code&amp;gt;a[i] &amp;lt;- sum(w[i,j] * x[j]) &amp;lt;/code&amp;gt;&lt;br /&gt;
## Iterate the previous step 1-5 times, your code should be able to (magically) restore the correct letter image even when the letter is mutated by 1-5 mutations in pixels. This exercise simulates the error-correction ability of memory (e.g., self-correcting encoding in CDs, a neon-light sign missing a stroke, or our ability to read/understand/reconstruct texts with typos, e.g., you have no problem reading/understanding this sentence: &amp;quot;It deosn’t mttaer in what order the ltteers in a word are, the olny iprmoatnt thing is taht the frist and lsat ltteer be in the rghit pclae.&amp;quot;). &lt;br /&gt;
## Expected capacity of a Hopfield Network: &amp;lt;code&amp;gt;number_of_memories = 0.138 * number_of_neurons&amp;lt;/code&amp;gt;. In our case, 25 (neurons) *0.138 = 3.45 memorized letters. So the network/memory fails if we squeeze in one additional letter (try it!).&lt;br /&gt;
# Code submissions&lt;br /&gt;
## Roy: [http://rpubs.com/roynunez/288433 rPubs Notebook]&lt;br /&gt;
## John: [https://github.com/JohnDi0505/MonteCarlo_Simulation-Biostats/blob/master/Neuron_Networks/Unsupervised_ML_Hopfield_Networks.ipynb Python Code on Github]&lt;br /&gt;
## Weigang: [http://rpubs.com/weigang/281806 rPubs Notebook]&lt;br /&gt;
# Potential biological applications: genetic code optimization; gene family identification&lt;br /&gt;
&lt;br /&gt;
==Week 7. Biological Applications==&lt;br /&gt;
A conceptual map for choosing ML algorithms: [[File:Ml map.png|thumbnail|Conceptual map from SciKit-Learn website]]&lt;br /&gt;
* The Python SciKit Learn framework may be tool of choice (instead of R). [http://scikit-learn.org/stable/auto_examples/index.html See these nice examples]  &lt;br /&gt;
* Bayesian Network (using e.g., R Package &amp;lt;code&amp;gt;bnlearn&amp;lt;/code&amp;gt;) to identify cell-signaling networks. Sache et al (2005). [http://science.sciencemag.org/content/308/5721/523.long Causal Protein-Signaling Networks Derived from Multiparameter Single-Cell Data]&lt;br /&gt;
* Identify proteins associated with learning in mice (Classification &amp;amp; Clustering)[https://archive.ics.uci.edu/ml/datasets/Mice+Protein+Expression UCI Mice Protein Expression Dataset]&lt;br /&gt;
* Predict HIV-1 protease cleavage sites (Classification): [https://archive.ics.uci.edu/ml/datasets/HIV-1+protease+cleavage UCI HIV-1 Protease Dataset]&lt;br /&gt;
* Lab project 1. Identify SNPs in 50 c-di-GMP pathway genes that are associated with c-di-GMP levels, swarming ability, and biofilm ability in 30 clinical isolates of Pseudomonas aeruginosa. Approach: supervised neural network (with regularization)&lt;br /&gt;
* Lab project 2. Identify genetic changes contributing to antibiotic resistance in 3 cancer patients. Approach: whole-genome sequencing followed by statistical analysis&lt;br /&gt;
* Lab project 3. Simulated evolution of genetic code. Approach: Multinomial optimization with unsupervised neural network&lt;br /&gt;
** An implementation example: [http://www.hoonzis.com/neural-networks-f-xor-classifier-and/ in C# language]&lt;br /&gt;
==Summer Project 1. Systems evolution of biofilm/swarming pathway (with Dr Joao Xavier of MSKCC)==&lt;br /&gt;
[[File:sim-cor-2.png|thumbnail| &amp;lt;b&amp;gt;Fig.1 .Simulated CDG-correlated SNPs.&amp;lt;/b&amp;gt; t-test results: (1) cor=0.8 (strong), t=-5.1142, df=95.596, p=1.619e-06; (2) cor=0.5 (medium), t=-4.7543, df=85.796, p=7.953e-06; (3) cor=0.2 (weak), t=-0.94585, df=79.28, p= 0.3471.]]&lt;br /&gt;
[[File:Sim-tree-snp-1.png|thumbnail|&amp;lt;b&amp;gt;Fig.2. Simulated tree-based SNPs.&amp;lt;/b&amp;gt; Generated with the APE function: &amp;lt;code&amp;gt;replicate(10, rTraitDisc(tr, states = c(0,1), rate = 100, model = &amp;quot;ER&amp;quot;))&amp;lt;/code&amp;gt;]]&lt;br /&gt;
# Acknowledgement: NSF Award 1517002&lt;br /&gt;
# Explanatory variable (genotypes): Whole-genome data (of ~30 clinical isolates &amp;amp; many experimentally evolved strains) as independent variables&lt;br /&gt;
# Explanatory variable (genotypes): ~50 genes related to cyclic-di-GMP synthesis and regulation (~8000 SNPs, ~1700 unique, ~5-8 major groups)&lt;br /&gt;
# Response variables (phenotypes): biofilm size, swarming size, antibiotic sensitivity profile, metabolomics measurements, c-di-GMP levels&lt;br /&gt;
# Sub-project 1: Database updates (Usmaan, Edgar, Christopher; continuing the work of Rayees and Raymond in previous years)&lt;br /&gt;
## c-di-GMP pathway SNPs in the new table &amp;quot;cdg_snp&amp;quot;&lt;br /&gt;
## c-di-GMP levels in &amp;quot;phenotype&amp;quot; table&lt;br /&gt;
## capture fig orthologs (in progress)&lt;br /&gt;
## capture matebolomics data with a new table (in process)&lt;br /&gt;
### [http://diverge.hunter.cuny.edu/~weigang/hm_metabolite-dendoclust.html a heatmap of deviation from mean (among strains of the same metabolite)]&lt;br /&gt;
### [http://diverge.hunter.cuny.edu/~weigang/matolites-fc-volcano-plot.html a volcano plot (p values based on t-test from group mean among strains of the same metabolite)]&lt;br /&gt;
# Sub-project 2. Supervised learning for predicting genes, SNPs associated with phenotypes&lt;br /&gt;
## Advantages over traditional regression analysis: ability to discover non-linear correction structure&lt;br /&gt;
## Challenges:&lt;br /&gt;
### over-fitting (we have only ~30 observations while thousands of SNPs, &amp;quot;curse of dimensionality&amp;quot;)&lt;br /&gt;
### the &amp;quot;effective&amp;quot; sample size is further discounted/reduced by phylogenetic relatedness among the strains.&lt;br /&gt;
## Single neuron, two-levels; by SNP groups (using &amp;lt;code&amp;gt;hclust()&amp;lt;/code&amp;gt;); by PCA (linear combination of SNPs); or by linkage groups (looking for homoplasy) (Mei), or by t-SNE (as suggested by Rayees)&lt;br /&gt;
## Three neurons, three-levels; by gene (Roy)[http://rpubs.com/roynunez/291340 july-14_nnet_cidigmp.R]&lt;br /&gt;
## To Do: (1) predict continuous-value targets; (2) add hidden layers; (3) correct for phylogenetic auto-correlation&lt;br /&gt;
# Sub-project 3. Generate simulated data (with Choleski Decomposition) (Weigang; continue the work in previous year by Ishmere, Zwar, &amp;amp; Rayees)&lt;br /&gt;
Fig.1. Simulated SNPs associated with cdg levels (w/o tree) [https://stats.stackexchange.com/questions/12857/generate-random-correlated-data-between-a-binary-and-a-continuous-variable inspired by this algorithm]&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
# simulate SNP-cdg correlation&lt;br /&gt;
r &amp;lt;- 0.2 # desired correlation coefficient&lt;br /&gt;
sigma &amp;lt;- matrix(c(1,r,r,1), ncol=2) # var-covariance matrix&lt;br /&gt;
s &amp;lt;- chol(sigma) # choleski decomposition&lt;br /&gt;
n &amp;lt;- 100 # number of random deviates (data points)&lt;br /&gt;
z &amp;lt;- s %*% matrix(rnorm(n*2), nrow=2) # 100 correlated normally distributed deviates with cor(x,y)=r&lt;br /&gt;
u &amp;lt;- pnorm(z) # get probabilities for each deviates&lt;br /&gt;
snp.states &amp;lt;- qbinom(u[1,], 1, 0.5) # discretize the 1st vector of probabilities into 0/1 with Bernoulli trial&lt;br /&gt;
idx.0 &amp;lt;- which(snp.states == 0); # indices for &amp;quot;0&amp;quot;&lt;br /&gt;
idx.1 &amp;lt;- which(snp.states == 1); # indices for &amp;quot;1&amp;quot;&lt;br /&gt;
# boxplots with stripcharts&lt;br /&gt;
boxplot(u[2,] ~ snp.states, main=&amp;quot;cor=0.2&amp;quot;, xlab=&amp;quot;SNP states&amp;quot;, ylab=&amp;quot;CDG level&amp;quot;)&lt;br /&gt;
stripchart(u[2,] ~ snp.states, vertical=T, pch=1, method=&amp;quot;jitter&amp;quot;, col=2,  add=T)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
Fig.2. Simulated SNPs associated with strain phylogeny&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
# Simulate SNPs on a tree&lt;br /&gt;
tr &amp;lt;- read.tree(&amp;quot;cdg-tree-mid.dnd&amp;quot;) # mid-point rooted tree&lt;br /&gt;
X &amp;lt;- replicate(10, rTraitDisc(tr, states = c(0,1), rate = 100, model = &amp;quot;ER&amp;quot;))&lt;br /&gt;
id &amp;lt;- read.table(&amp;quot;cdg.strains.txt3&amp;quot;, row.names = 1, sep=&amp;quot;\t&amp;quot;)&lt;br /&gt;
par.tr &amp;lt;- plot(tr, no.margin = T, x.lim = 0.03, show.tip.label = F)&lt;br /&gt;
text(rep(0.013,30), 1:30, id[tr$tip.label,1], pos = 4, cex=0.75)&lt;br /&gt;
text(rep(0.02,30), 1:30, snps[tr$tip.label], pos=4, cex=0.75)&lt;br /&gt;
add.scale.bar()&lt;br /&gt;
abline(h=1:30, col=&amp;quot;gray&amp;quot;, lty=2)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
To Do/Challenges: (1) simulate multiple SNPs (linear correlation); (2) simulate epistasis (non-linear correlation); (3) simulate phylogenetic correlation&lt;br /&gt;
* Simulation &amp;amp; xgboost code contributed by Mei &amp;amp; Yinheng (January 2018)&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
&lt;br /&gt;
#To create a simulated matrix w/ correlated xy variables; NOTE: ONLY 1 correlated x variable&lt;br /&gt;
get1Simulated &amp;lt;- function(corco, nstrains, snps){&lt;br /&gt;
  r &amp;lt;- corco/10 # desired correlation coefficient&lt;br /&gt;
  sigma &amp;lt;- matrix(c(1,r,r,1), ncol=2) # var-covariance matrix&lt;br /&gt;
  s &amp;lt;- chol(sigma) #cholesky decomposition&lt;br /&gt;
  n &amp;lt;- nstrains&lt;br /&gt;
  z &amp;lt;- s %*% matrix(rnorm(n*2), nrow=2)&lt;br /&gt;
  u &amp;lt;- pnorm(z[2,])&lt;br /&gt;
  snp.states &amp;lt;- qbinom(u, 1, 0.5)&lt;br /&gt;
&lt;br /&gt;
  known &amp;lt;- t(rbind(z[1,],snp.states))&lt;br /&gt;
  rand &amp;lt;- pnorm(matrix(rnorm(nstrains*(snps-1)),nrow = nstrains))&lt;br /&gt;
  snp.rand&amp;lt;-qbinom(rand,1,0.5)&lt;br /&gt;
  x &amp;lt;- cbind(known, snp.rand)&lt;br /&gt;
&lt;br /&gt;
  colnames(x)[1] &amp;lt;- &amp;quot;target&amp;quot;&lt;br /&gt;
  colnames(x)[-1] &amp;lt;- paste(&amp;quot;feature&amp;quot;, seq_len(ncol(x)-1), sep = &amp;quot;.&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
  return(x)&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
#Use this function to artificially create 2 x variables that have relationship w/ Y.&lt;br /&gt;
getSimulated.chol &amp;lt;- function(snp1, snp2, nstrains, snps){&lt;br /&gt;
  n &amp;lt;- nstrains&lt;br /&gt;
&lt;br /&gt;
  btwn &amp;lt;- snp1 * snp2 #this is a suggested value for inter-SNP correlation. the covariance matrix has to be semi-positive infinite in order to carry out the decomposition&lt;br /&gt;
&lt;br /&gt;
  covar.m &amp;lt;- matrix(c(1.0, snp1, snp2, snp1, 1.0, btwn, snp2, btwn, 1.0), nrow = 3) #cor-matrix&lt;br /&gt;
&lt;br /&gt;
  s &amp;lt;- chol(covar.m)&lt;br /&gt;
  z &amp;lt;- s %*% matrix(rnorm(n*3), nrow = 3)&lt;br /&gt;
&lt;br /&gt;
  #decretize the 2 variables&lt;br /&gt;
  u &amp;lt;- t(pnorm(z[2:nrow(z),]))&lt;br /&gt;
  decretize.u &amp;lt;-qbinom(u,1,0.5)&lt;br /&gt;
&lt;br /&gt;
  #generate some random discrete variables and combine w/ the 2 correlated variables&lt;br /&gt;
  rand &amp;lt;- matrix(rnorm(nstrains*(snps-2)),nrow = nstrains)&lt;br /&gt;
  rand.u &amp;lt;- pnorm(rand)&lt;br /&gt;
  snp.rand &amp;lt;- qbinom(rand.u,1,0.5)&lt;br /&gt;
  x &amp;lt;- cbind(z[1,], decretize.u, snp.rand)&lt;br /&gt;
&lt;br /&gt;
  dimnames(x) &amp;lt;- list(c(), c(&amp;quot;target&amp;quot;, paste(&amp;quot;feature&amp;quot;, seq_len(ncol(x)-1), sep = &amp;quot;.&amp;quot;)))&lt;br /&gt;
&lt;br /&gt;
  return(x)&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
#This utilizes xgboost function w/o cross-validating the strains and w/ tuned parameters (we think it&#039;s the optimal set) tested by Yinheng&#039;s python [https://github.com/weigangq/mic-boost/blob/master/micboost/__init__.py getBestParameters] function. Please make sure to have xgboost installed. &lt;br /&gt;
runXg &amp;lt;- function(x){&lt;br /&gt;
  require(xgboost)&lt;br /&gt;
&lt;br /&gt;
  bst &amp;lt;- xgboost(data = x[,2:ncol(x)], label = x[,1], max.depth = 2, eta = .05, gamma = 0.3, nthread = 2, nround = 10, verbose = 0, eval_metric = &amp;quot;rmse&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
  importance_matrix &amp;lt;- xgb.importance(model = bst, feature_names = colnames(x[,2:ncol(x)]))&lt;br /&gt;
&lt;br /&gt;
  return(importance_matrix)&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
#Same as above but w/ cross validation.&lt;br /&gt;
runXg.cv &amp;lt;- function(x){&lt;br /&gt;
  require(caret)&lt;br /&gt;
  require(xgboost)&lt;br /&gt;
&lt;br /&gt;
  ind &amp;lt;- createDataPartition(x[,1], p = 2/3, list = FALSE )&lt;br /&gt;
&lt;br /&gt;
  trainDf &amp;lt;- as.matrix(x[ind,])&lt;br /&gt;
  testDf &amp;lt;- as.matrix(x[-ind,])&lt;br /&gt;
&lt;br /&gt;
  dtrain &amp;lt;- xgb.DMatrix(data = trainDf[,2:ncol(trainDf)], label = trainDf[,1])&lt;br /&gt;
  dtest &amp;lt;- xgb.DMatrix(data = testDf[,2:ncol(testDf)], label = testDf[,1])&lt;br /&gt;
&lt;br /&gt;
  watchlist &amp;lt;- list(train=dtrain, test=dtest)&lt;br /&gt;
&lt;br /&gt;
  bst &amp;lt;- xgb.train(data=dtrain, nthread = 2, nround=10, watchlist=watchlist, eval.metric = &amp;quot;rmse&amp;quot;, verbose = 0)&lt;br /&gt;
&lt;br /&gt;
  importance_matrix &amp;lt;- xgb.importance(model = bst, feature_names = colnames(x[,2:ncol(x)]))&lt;br /&gt;
&lt;br /&gt;
  return(importance_matrix)&lt;br /&gt;
&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
#This function returns a master table with different iterated values that finds the best number of predictor variables (x).&lt;br /&gt;
getTable &amp;lt;- function(x, importance_matrix, cor){&lt;br /&gt;
  # v.names &amp;lt;- c(&amp;quot;feature&amp;quot;, &amp;quot;SNPs&amp;quot;, &amp;quot;strains&amp;quot;, &amp;quot;cor_given&amp;quot;, &amp;quot;cor_actual&amp;quot;, &amp;quot;gain&amp;quot;, &amp;quot;cover&amp;quot;, &amp;quot;rank&amp;quot;)&lt;br /&gt;
  v.names &amp;lt;- c(&amp;quot;feature&amp;quot;, &amp;quot;SNPs&amp;quot;, &amp;quot;strains&amp;quot;, &amp;quot;cor_given&amp;quot;, &amp;quot;cor_actual&amp;quot;, &amp;quot;p.value&amp;quot;, &amp;quot;gain&amp;quot;, &amp;quot;cover&amp;quot;, &amp;quot;rank&amp;quot;)&lt;br /&gt;
  for (i in 1:length(v.names)){&lt;br /&gt;
    assign(v.names[i], numeric())&lt;br /&gt;
  }&lt;br /&gt;
  featNum &amp;lt;- NULL&lt;br /&gt;
&lt;br /&gt;
  n &amp;lt;- 0&lt;br /&gt;
  new_cor &amp;lt;- rep(cor, length(x)/length(cor))&lt;br /&gt;
  for (num in 1:length(x)) {&lt;br /&gt;
    featNum &amp;lt;- NULL&lt;br /&gt;
    imp_matrix &amp;lt;- importance_matrix[[num]]&lt;br /&gt;
    x.iter &amp;lt;- x[[num]]&lt;br /&gt;
    x.pv &amp;lt;- getPV(x[[num]])&lt;br /&gt;
&lt;br /&gt;
    for (i in 1:2){&lt;br /&gt;
      featNum &amp;lt;- c(featNum, grep(paste(&amp;quot;feature.&amp;quot;,i,&amp;quot;$&amp;quot;, sep = &amp;quot;&amp;quot;), imp_matrix$Feature, perl = TRUE))&lt;br /&gt;
    }&lt;br /&gt;
&lt;br /&gt;
    n &amp;lt;-  n + 1&lt;br /&gt;
&lt;br /&gt;
    for(i in 1:2){&lt;br /&gt;
      f &amp;lt;- featNum[i]&lt;br /&gt;
      feature&amp;lt;-c(feature, imp_matrix$Feature[f])&lt;br /&gt;
      SNPs &amp;lt;-c(SNPs, 100)&lt;br /&gt;
      strains &amp;lt;- c(strains, nrow(x.iter))&lt;br /&gt;
      cor_given &amp;lt;- c(cor_given, new_cor[n])&lt;br /&gt;
      cor_actual &amp;lt;-c(cor_actual, cor(x.iter[,1],x.iter[,grep(paste(imp_matrix$Feature[f], &amp;quot;$&amp;quot;, sep = &amp;quot;&amp;quot;), colnames(x.iter), perl = TRUE)]))&lt;br /&gt;
      p.value &amp;lt;- c(p.value, x.pv[,imp_matrix$Feature[f]])&lt;br /&gt;
      gain &amp;lt;- c(gain, imp_matrix$Gain[f])&lt;br /&gt;
      cover &amp;lt;- c(cover, imp_matrix$Cover[f])&lt;br /&gt;
      rank &amp;lt;- c(rank, f)&lt;br /&gt;
    }&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
  x.master &amp;lt;- data.frame(feature, SNPs, strains, cor_given, cor_actual, p.value, gain, cover, rank)&lt;br /&gt;
&lt;br /&gt;
  return(x.master)&lt;br /&gt;
}&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Summer Project 2. Whole-genome variants associated with multidrug resistance in clinical Pseudomonas &amp;amp; E.coli isolates (with Dr Yi-Wei Tang of MSKCC)==&lt;br /&gt;
# Acknowledgement: Hunter CTBR Pilot Award&lt;br /&gt;
# Stage 1: Select patients &amp;amp; strains for whole-genome sequencing by MiSeq, based on drug sensitivities (16 isolates from 5 patients were isolated, tested, and selected by April, 2017)&lt;br /&gt;
# Stage 2: Genome sequencing (FASTQ files generated, 3 replicates for each isolates; done by June 2017)&lt;br /&gt;
# Stage 3: Variant call&lt;br /&gt;
## Reference strains identified using Kraken (Roy)&lt;br /&gt;
## VCF generation using cortex_var (Michele and Hanna, led by John)&lt;br /&gt;
# Stage 4: Variant annotation&lt;br /&gt;
# Stage 5: Statistical analysis&lt;br /&gt;
# Stage 6: Web report&lt;br /&gt;
&lt;br /&gt;
==Summer Project 3. Origin of Genetic Code==&lt;br /&gt;
Tools for testing SGC &amp;amp; evolved codes&lt;br /&gt;
* Shuffle code: &lt;br /&gt;
&amp;lt;syntaxhighlight lang=&#039;bash&#039;&amp;gt;&lt;br /&gt;
./shuffle-code.pl&lt;br /&gt;
    -f &amp;lt;sgc|code-file&amp;gt; (required)&lt;br /&gt;
    -s &amp;lt;1|2|3|4|5&amp;gt; (shuffle by 1st, 2nd, 3rd, aa blocks, and all random)&lt;br /&gt;
    -p(olarity; default)&lt;br /&gt;
    -h(ydropathy)&lt;br /&gt;
    -v(olume)&lt;br /&gt;
    -e(iso-electricity)&lt;br /&gt;
# Input: &#039;sgc&#039; (built-in) or an evolved code file consisting of 64 rows of &amp;quot;codon&amp;quot;-&amp;quot;position&amp;quot;&lt;br /&gt;
# Output: a code file consisting of 64 rows of &amp;quot;codon&amp;quot; - &amp;quot;aa&amp;quot; (to be fed into ./code-stats.pl)&lt;br /&gt;
# Usage examples:&lt;br /&gt;
./shuffle-code.pl -f &#039;sgc&#039; -s 5 # randomly permute SGC&lt;br /&gt;
./shuffle-code.pl -f &#039;evolved-code.txt&#039; -p # evolved code, AA assigned according to polarity (default)&lt;br /&gt;
./shuffle-code.pl -f &#039;evolved-code.txt&#039; -s 5 -h # evolved code, AA assigned according to hydrophobicity, shuffled randomly&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Code statistics&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&#039;bash&#039;&amp;gt;&lt;br /&gt;
./code-stats.pl&lt;br /&gt;
        [-h]    (help)&lt;br /&gt;
        [-f &#039;sgc (default)|code_file&#039;]  (code)&lt;br /&gt;
        [-s &#039;pair (default)|fit|path&#039;]  (stats)&lt;br /&gt;
        [-p &#039;pol|hydro|vol|iso&#039;, default &#039;grantham&#039;]    (aa prop)&lt;br /&gt;
        [-i (ti/tv, default 5)]&lt;br /&gt;
        [-b: begin codon (&#039;TTT&#039;) -q: panelty for 1st (50); -w: panelty for 2nd (100)]   (options for path)&lt;br /&gt;
# Input: &#039;sgc&#039; (built-in) or a code file consisting of 64 rows of &amp;quot;codon&amp;quot; - &amp;quot;aa&amp;quot;&lt;br /&gt;
# Output: codon or code statistics&lt;br /&gt;
# Usage examples:&lt;br /&gt;
./code-stats.pl # all defaults: print single-mutation codon pairs, grantham distance, for SGC&lt;br /&gt;
./code-stats.pl -s &#039;fit&#039; -p &#039;pol&#039; # print code fitness according to polarity, for SGC (default)&lt;br /&gt;
./code-stats.pl -s &#039;path&#039; # print tour length for SGC&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;gallery perrow=&amp;quot;4&amp;quot;&amp;gt;&lt;br /&gt;
Sgc-polar.png|SGC: ordered by polarity&lt;br /&gt;
Sgc-hydro.png|SGC: ordered by hydropathy&lt;br /&gt;
Sgc-vol.png|SGC: ordered by volume&lt;br /&gt;
Sgc-iso.png|SGC: ordered by isoelectricity&lt;br /&gt;
&lt;br /&gt;
Path-code3.png|An evolved code (by Oliver)&lt;br /&gt;
code-john-2.png|An evolved code (by John). Round-trip (last position connected with the first) AA assigned according to polarity gradient &amp;amp; with &amp;quot;TAA&amp;quot; as the 1st.&lt;br /&gt;
code-brian-1.png|An evolved code (by Brian). One-way tour (last position NOT connected with the first). AA assigned according to polarity gradient &amp;amp; with &amp;quot;TAA&amp;quot; as the 1st.&lt;br /&gt;
code-error-v3.png|code robustness: &amp;lt;font color=&amp;quot;red&amp;quot;&amp;gt;SGC&amp;lt;/font&amp;gt;, &amp;lt;font color=&amp;quot;blue&amp;quot;&amp;gt;John&#039;s simulated code &amp;lt;/font&amp;gt;, &amp;lt;font color=&amp;quot;orange&amp;quot;&amp;gt;Brian&#039;s simulated code &amp;lt;/font&amp;gt;, histogram: 1000 random codes. X-axis: mean-squared error (standardized) caused by single-nt substitutions. Evolved codes perform better than SGC!&lt;br /&gt;
&lt;br /&gt;
Aa.pca.png|PC1 (mostly polarity and hydropathy, anti-correlated) &amp;amp; PC2 (the other two, positively correlated) of 4 AA metrics. PC1 and PC2 could be used for AA ranking as composite variables&lt;br /&gt;
tRNA-tree.png|A tRNA gene tree from Aeropyrum pernix (an Archaea). Sequences from [http://trna.bioinf.uni-leipzig.de/DataOutput/ an rRNA database]. Structual alignment available&lt;br /&gt;
tRNA-tree-2.png|tRNA gene tree for Pyrococcus horikoshii (another Archaea). Sequences from [http://gtrnadb.ucsc.edu/GtRNAdb2/index.html UCSD tRNA database]. (Not structurally aligned; only fasta seqs available)&lt;br /&gt;
tRNA-seq-dist.png|AA distance appears to be correlated with tRNA seq differences (especially at low distances)&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
# Participants: Oliver, John, Brian&lt;br /&gt;
# A representation mimicking TSP (see figure at right)&lt;br /&gt;
# Code to calculate total (mutational, ti/tv) path of SGC?&lt;br /&gt;
# Code to calculate energy of SGC?&lt;br /&gt;
# Randomize amino acid assignment and obtain distributions of path &amp;amp; energy&lt;br /&gt;
# Code to minimize total path length and energy?&lt;br /&gt;
# How to evolve a robust SGC? mutation bias + polarity/hydropathy + usage&lt;br /&gt;
&lt;br /&gt;
=Season I. Spring 2017 (Themes: Dueling Idiots/Digital Dice/Wright Fisher Process)=&lt;br /&gt;
==&amp;quot;Coalescence&amp;quot; (Backward simulation of Wright-Fisher process) (Due May 12, 2017)==&lt;br /&gt;
[[File:Coalescent-tree.png|thumbnail|([http://raven.iab.alaska.edu/~ntakebay/teaching/programming/coalsim/node1.html Source])]]&lt;br /&gt;
[[File:Coalescent-output-1.png|thumbnail]]&lt;br /&gt;
We will conclude Spring 2017 season with the coalescence simulation (in summer, we will start experimenting with simulation of systems evolution) &lt;br /&gt;
&lt;br /&gt;
Previously, we simulated Wright-Fisher process of genetic drift (constant pop size, no selection) starting from a founder population and end with the present population. This is not efficient because the program has to track each individual in each generation, although the majority of them do not contribute to the present population.&lt;br /&gt;
&lt;br /&gt;
Coalescence simulation takes the opposite approach of starting from the present sample of k individuals and trace backward in time to their most recent common ancestor (MRCA). Due to the nature of Poisson process, the waiting time from one coalescence event (at time T) to the next one (at time T-1) is exponentially distributed with a mean of (k choose 2) generations. This process is iterated until the last coalescent event.&lt;br /&gt;
&lt;br /&gt;
The classic text on coalescence is [http://home.uchicago.edu/rhudson1/popgen356/OxfordSurveysEvolBiol7_1-44.pdf Richard Hudson&#039;s chapter], which includes  (in Appendix) C codes for simulating tree and mutations. Hudson is also the author of &amp;lt;code&amp;gt;ms&amp;lt;/code&amp;gt;, the widely used coalescence simulator.&lt;br /&gt;
&lt;br /&gt;
Like the &amp;lt;code&amp;gt;ms 10 1 -T&amp;lt;/code&amp;gt; command, your code should output a tree of 10 individuals. [A lot tougher than I thought; can&#039;t figure it out in R; so I did it in Perl]. Similarly, you could use R, where the &amp;lt;code&amp;gt;ape&amp;lt;/code&amp;gt; package has a function called &amp;lt;code&amp;gt;rcoal()&amp;lt;/code&amp;gt; that generate a random coalescence tree. Try this command: &amp;lt;code&amp;gt;plot(rcoal(10))&amp;lt;/code&amp;gt; (first load library by running &amp;lt;code&amp;gt;library(ape)&amp;lt;/code&amp;gt;).&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
|&lt;br /&gt;
*By Weigang (First draft)&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;perl&amp;quot;&amp;gt;&lt;br /&gt;
#!/usr/bin/env perl&lt;br /&gt;
# Basic coalescence&lt;br /&gt;
# First, simulate coalescent events with recursion&lt;br /&gt;
# Second, use Bio::Tree to export a newick tree&lt;br /&gt;
use strict;&lt;br /&gt;
use warnings;&lt;br /&gt;
use Data::Dumper;&lt;br /&gt;
use Algorithm::Numerical::Sample  qw /sample/;&lt;br /&gt;
use Math::Random qw(random_exponential);&lt;br /&gt;
use Bio::Tree::Tree;&lt;br /&gt;
use Bio::Tree::Node;&lt;br /&gt;
&lt;br /&gt;
######################&lt;br /&gt;
# Initialize&lt;br /&gt;
######################&lt;br /&gt;
die &amp;quot;Usage: $0 &amp;lt;num-of-samples&amp;gt;\n&amp;quot; unless @ARGV == 1;&lt;br /&gt;
my $nsamp = shift @ARGV;&lt;br /&gt;
my @samples;&lt;br /&gt;
for (my $i=1; $i&amp;lt;=$nsamp; $i++) { push @samples, {id=&amp;gt;$i, parent=&amp;gt;undef, br=&amp;gt;0} }&lt;br /&gt;
my $ctr = $nsamp;&lt;br /&gt;
my $time = 0;&lt;br /&gt;
my @all_nodes;&lt;br /&gt;
&lt;br /&gt;
################################################################################&lt;br /&gt;
# Simulate events with exponential time intervals between two successive events&lt;br /&gt;
#################################################################################&lt;br /&gt;
&amp;amp;coal(\@samples, \$ctr, \$time);&lt;br /&gt;
#print Dumper(\@all_nodes);&lt;br /&gt;
#########################################&lt;br /&gt;
# Reconstitute into tree using Bio::Tree&lt;br /&gt;
##########################################&lt;br /&gt;
my %seen_node;&lt;br /&gt;
my $max_id=0;&lt;br /&gt;
my @nodes;&lt;br /&gt;
foreach (@all_nodes) {&lt;br /&gt;
    $max_id = ($_-&amp;gt;{parent} &amp;gt; $max_id) ? $_-&amp;gt;{parent} : $max_id;&lt;br /&gt;
    if ($seen_node{$_-&amp;gt;{id}}) { # child exist, previously as a parent, no branch length&lt;br /&gt;
        $seen_node{$_-&amp;gt;{id}}-&amp;gt;branch_length(sprintf &amp;quot;%.6f&amp;quot;, $_-&amp;gt;{br}); # add branch length&lt;br /&gt;
        if ($seen_node{$_-&amp;gt;{parent}}) { # parent exist&lt;br /&gt;
            $seen_node{$_-&amp;gt;{parent}}-&amp;gt;add_Descendent($seen_node{$_-&amp;gt;{id}});&lt;br /&gt;
        } else { # parent new&lt;br /&gt;
            my $pa = Bio::Tree::Node-&amp;gt;new(-id=&amp;gt;$_-&amp;gt;{parent});&lt;br /&gt;
            $pa-&amp;gt;add_Descendent($seen_node{$_-&amp;gt;{id}});&lt;br /&gt;
            $seen_node{$_-&amp;gt;{parent}} = $pa;&lt;br /&gt;
        }&lt;br /&gt;
    } else { # child new&lt;br /&gt;
        my $new = Bio::Tree::Node-&amp;gt;new(-id=&amp;gt;$_-&amp;gt;{id}, -branch_length =&amp;gt; sprintf &amp;quot;%.6f&amp;quot;, $_-&amp;gt;{br});&lt;br /&gt;
        $seen_node{$_-&amp;gt;{id}} = $new;&lt;br /&gt;
        if ($seen_node{$_-&amp;gt;{parent}}) { # parent exist&lt;br /&gt;
            $seen_node{$_-&amp;gt;{parent}}-&amp;gt;add_Descendent($new);&lt;br /&gt;
        } else { # parent new&lt;br /&gt;
            my $pa = Bio::Tree::Node-&amp;gt;new(-id=&amp;gt;$_-&amp;gt;{parent});&lt;br /&gt;
            $pa-&amp;gt;add_Descendent($new);&lt;br /&gt;
            $seen_node{$_-&amp;gt;{parent}} = $pa;&lt;br /&gt;
        }&lt;br /&gt;
    }&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
#my $root = $seen_node{$max_id};&lt;br /&gt;
my $tree=Bio::Tree::Tree-&amp;gt;new(-id=&amp;gt;&#039;coal-sim&#039;, -node=&amp;gt;$seen_node{$max_id}, -nodelete=&amp;gt;1);&lt;br /&gt;
print $tree-&amp;gt;as_text(&amp;quot;newick&amp;quot;), &amp;quot;\n&amp;quot;;&lt;br /&gt;
&lt;br /&gt;
exit;&lt;br /&gt;
&lt;br /&gt;
sub coal {&lt;br /&gt;
    my $ref_nodes = shift;&lt;br /&gt;
    my $ref_ct = shift;&lt;br /&gt;
    my $ref_time = shift;&lt;br /&gt;
    my $ct = $$ref_ct;&lt;br /&gt;
    my @current_nodes = @$ref_nodes;&lt;br /&gt;
    my $k = scalar @current_nodes;&lt;br /&gt;
    my @new_nodes;&lt;br /&gt;
    return unless $k &amp;gt; 1;&lt;br /&gt;
    my @pair = sample(-set =&amp;gt; $ref_nodes, -sample_size =&amp;gt; 2);&lt;br /&gt;
#    print $ct, &amp;quot;\t&amp;quot;, $$ref_time, &amp;quot;\t&amp;quot;, $pair[0]-&amp;gt;{id}, &amp;quot;\t&amp;quot;, $pair[1]-&amp;gt;{id}, &amp;quot;\n&amp;quot;;&lt;br /&gt;
    $$ref_time += random_exponential(1, 2/$k/($k-1));&lt;br /&gt;
    my $new_nd = {id=&amp;gt;$ct+1, parent=&amp;gt;undef, br=&amp;gt;$$ref_time};&lt;br /&gt;
    map {$_-&amp;gt;{parent} = $ct+1} @pair;&lt;br /&gt;
    map {$_-&amp;gt;{br} = $$ref_time - $_-&amp;gt;{br}} @pair;&lt;br /&gt;
    push @all_nodes, $_ for @pair;&lt;br /&gt;
    foreach (@current_nodes) {&lt;br /&gt;
        push @new_nodes, $_ unless $_-&amp;gt;{id} == $pair[0]-&amp;gt;{id} || $_-&amp;gt;{id} == $pair[1]-&amp;gt;{id};&lt;br /&gt;
    }&lt;br /&gt;
    push @new_nodes, $new_nd;&lt;br /&gt;
    $$ref_ct++;&lt;br /&gt;
    &amp;amp;coal(\@new_nodes, $ref_ct, $ref_time);&lt;br /&gt;
}&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
*By John&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;python&amp;quot;&amp;gt;&lt;br /&gt;
from Bio import Phylo&lt;br /&gt;
from io import StringIO&lt;br /&gt;
from random import sample&lt;br /&gt;
from scipy.misc import comb&lt;br /&gt;
from itertools import combinations&lt;br /&gt;
from numpy.random import exponential as exp&lt;br /&gt;
&lt;br /&gt;
# Generate Random Tree with Nodes &amp;amp; Waiting Time&lt;br /&gt;
individuals = [&amp;quot;A&amp;quot;, &amp;quot;B&amp;quot;, &amp;quot;C&amp;quot;, &amp;quot;D&amp;quot;, &amp;quot;E&amp;quot;, &amp;quot;F&amp;quot;, &amp;quot;G&amp;quot;, &amp;quot;H&amp;quot;, &amp;quot;I&amp;quot;, &amp;quot;J&amp;quot;]&lt;br /&gt;
node_dict = {}&lt;br /&gt;
reference = {}&lt;br /&gt;
wait_time_ref = {}&lt;br /&gt;
node = 0&lt;br /&gt;
while len(individuals) != 1:&lt;br /&gt;
    node_dict[node] = {&amp;quot;child_nodes&amp;quot;: []}&lt;br /&gt;
    sample_events = tuple(sample(individuals, 2))&lt;br /&gt;
    reference[sample_events] = node&lt;br /&gt;
    wait_time = exp(1/comb(len(individuals), 2))&lt;br /&gt;
    wait_time_ref[node] = wait_time&lt;br /&gt;
    for sample_event in sample_events:&lt;br /&gt;
        if type(sample_event) == str:&lt;br /&gt;
            node_dict[node][&amp;quot;child_nodes&amp;quot;].append(sample_event)&lt;br /&gt;
        else:&lt;br /&gt;
            child_node = reference[sample_event]&lt;br /&gt;
            node_dict[node][&amp;quot;child_nodes&amp;quot;].append(child_node)&lt;br /&gt;
    individuals.remove(sample_events[0])&lt;br /&gt;
    individuals.remove(sample_events[1])&lt;br /&gt;
    individuals.append(sample_events)&lt;br /&gt;
    node += 1&lt;br /&gt;
    &lt;br /&gt;
# Calculate Tree Branch Lengths&lt;br /&gt;
tree_dict = {}&lt;br /&gt;
cumulative_br_length = 0&lt;br /&gt;
cumulative_br_len_dict = {}&lt;br /&gt;
for i in range(len(node_dict)):&lt;br /&gt;
    tree_dict[i] = {}&lt;br /&gt;
    cumulative_br_length += wait_time_ref[i]&lt;br /&gt;
    cumulative_br_len_dict[i] = cumulative_br_length&lt;br /&gt;
    for node in node_dict[i][&#039;child_nodes&#039;]:&lt;br /&gt;
        if type(node) == str:&lt;br /&gt;
            tree_dict[i][node] = cumulative_br_length&lt;br /&gt;
        else:&lt;br /&gt;
            tree_dict[i][node] = cumulative_br_len_dict[i] - cumulative_br_len_dict[node]&lt;br /&gt;
            &lt;br /&gt;
# Parse the Tree into a String&lt;br /&gt;
for i in range(len(tree_dict)):&lt;br /&gt;
    for node in tree_dict[i].keys():&lt;br /&gt;
        if type(node) != str:&lt;br /&gt;
            temp = str(tree_dict[node])&lt;br /&gt;
            tree_dict[i][temp] = tree_dict[i].pop(node)&lt;br /&gt;
tree_str = str(tree_dict[i]).replace(&#039;\&#039;&#039;, &#039;&#039;).replace(&#039;\&amp;quot;&#039;, &#039;&#039;).replace(&#039;\\&#039;, &#039;&#039;).replace(&#039;{&#039;, &#039;(&#039;).replace(&#039;}&#039;, &#039;)&#039;).replace(&#039; &#039;, &#039;&#039;)&lt;br /&gt;
tree_str += &amp;quot;;&amp;quot;            &lt;br /&gt;
&lt;br /&gt;
# Visualize Tree&lt;br /&gt;
handle = StringIO(tree_str)&lt;br /&gt;
tree = Phylo.read(handle, &#039;newick&#039;)&lt;br /&gt;
tree.ladderize()   # Flip branches so deeper clades are displayed at top&lt;br /&gt;
Phylo.draw(tree)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==April 17, 2017 &amp;quot;Mutation Meltdown&amp;quot; (Due May 5, 2017)==&lt;br /&gt;
[[File:Ratchet-1.png|thumbnail]]&lt;br /&gt;
Our last exploration showed that population will reach a steady-state level of DNA sequence polymorphism under the opposing forces of genetic drift and mutations. The steady-state level is expected to be ~ N * mu: the larger the population size and the higher the mutation rate, the higher per-site DNA polymorphism. (Also note that, given a long enough sequence and a low enough mutation rate, we don&#039;t expect to see two or more mutations hitting a single position. Each mutation is essentially a new one and only two-state SNPs are expected in a sample of DNA sequences).&lt;br /&gt;
&lt;br /&gt;
However, the steady-state expectation is based on the assumption that all mutations are neutral (i.e., no beneficial or harmful fitness effects). In reality, mutations are predominantly either neutral or harmful and few are beneficial. Natural selection (negative or positive, except those on the immune-defense loci) tends to drive down genetic variation.&lt;br /&gt;
&lt;br /&gt;
Sex to the rescue. Without sex or recombination, the fate of genetic variations across a genome are bundled together, rising or sinking in unison with the fate of a single beneficial or harmful mutation. With recombination, genetic variations at different loci become less tightly linked and are more likely maintained, speeding up adaptation.&lt;br /&gt;
&lt;br /&gt;
This week, we will use simulation to recreate the so-called &amp;quot;[https://en.wikipedia.org/wiki/Muller%27s_ratchet Muller&#039;s Ratchet]&amp;quot;, which predicts that asexual populations are evolutionary dead ends. It shows that the combined effect of deleterious mutation and genetic drift will lead to population extinction due to steady accumulation of deleterious mutations, a phenomenon called &amp;quot;mutation meltdown&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
The simulation protocol would be similar to that of the last problem involving only mutation and drift, but I suggest you rewrite from scratch to be more efficient by using a simpler data structure (no sequence or bases needed). The main difference is that, instead of calculating average pairwise sequence differences in each generation, you will track the frequency of mutation-free sequences (n0) and find the time until it goes to zero. That is when the Ratchet makes a &amp;quot;click&amp;quot;. The next click is when the population losses the one-mutation sequences (n1), and so on. Each click drives the population fitness down by one mutation-level. [ Numerically, if each mutation causes a fitness loss of &amp;lt;code&amp;gt;s&amp;lt;/code&amp;gt; (&amp;quot;selection coefficient&amp;quot;), the fitness of a sequence with &amp;lt;code&amp;gt;k&amp;lt;/code&amp;gt; mutations is given by w = (1-s)&amp;lt;sup&amp;gt;k&amp;lt;/sup&amp;gt;. Strictly speaking, the selection coefficient should be included to affect gamete size, but let&#039;s ignore that for now].&lt;br /&gt;
&lt;br /&gt;
If there is recombination, the mutation-free sequences could be recovered (simulation next time?). Without recombination, the only direction for the population to evolve is a steady loss of best-fit individuals (by genetic drift).&lt;br /&gt;
&lt;br /&gt;
Questions:&lt;br /&gt;
# Does the ratchet occur faster in small or large populations? Find by simulating N=100 and N=1000&lt;br /&gt;
# Does the ratchet occur faster for a short or long genome? Find by simulating L=1e3 and L=1e4&lt;br /&gt;
# Which parts of the human genomes are asexual (therefore subject to Muller&#039;s Ratchet and becoming increasingly small)?&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
|&lt;br /&gt;
*By John&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;python&amp;quot;&amp;gt;&lt;br /&gt;
import numpy as np&lt;br /&gt;
import matplotlib.pyplot as plt&lt;br /&gt;
from numpy.random import poisson&lt;br /&gt;
from numpy.random import choice&lt;br /&gt;
&lt;br /&gt;
def simulator(seq_length, pop, repro_rate, generations):&lt;br /&gt;
    mutation_rate = 0.00001&lt;br /&gt;
    lam = seq_length * mutation_rate&lt;br /&gt;
    individuals = np.array([1 for i in range(pop)])&lt;br /&gt;
    non_mutated_pop_rate = []&lt;br /&gt;
    for generation in range(generations):&lt;br /&gt;
        gametes = []&lt;br /&gt;
        for individual in individuals:&lt;br /&gt;
            off_sprs = poisson(lam, repro_rate)&lt;br /&gt;
            mutation_ix = np.array(np.where(off_sprs != 0))&lt;br /&gt;
            indiv_gametes = np.array([individual for j in range(repro_rate)])&lt;br /&gt;
            indiv_gametes[mutation_ix] = 0&lt;br /&gt;
            gametes += list(indiv_gametes)&lt;br /&gt;
        individuals = choice(gametes, size=pop)&lt;br /&gt;
        non_mutated_frequency = float(np.count_nonzero(individuals)) / pop&lt;br /&gt;
        non_mutated_pop_rate.append(non_mutated_frequency)&lt;br /&gt;
    return non_mutated_pop_rate&lt;br /&gt;
&lt;br /&gt;
genome_length_1000 = 1000&lt;br /&gt;
genome_length_10000 = 10000&lt;br /&gt;
results_1000 = simulator(genome_length_1000, 1000, 100, 500)&lt;br /&gt;
results_10000 = simulator(genome_length_10000, 1000, 100, 500)&lt;br /&gt;
&lt;br /&gt;
genome_length_1000 = 1000&lt;br /&gt;
genome_length_10000 = 10000&lt;br /&gt;
results2_1000 = simulator(genome_length_1000, 10000, 100, 500)&lt;br /&gt;
results2_10000 = simulator(genome_length_10000, 10000, 100, 500)&lt;br /&gt;
&lt;br /&gt;
f, (ax1, ax2) = plt.subplots(1, 2, sharey=True)&lt;br /&gt;
line_1_1000 = ax1.plot(results_1000, &amp;quot;r&amp;quot;)&lt;br /&gt;
line_1_10000 = ax1.plot(results_10000, &amp;quot;b&amp;quot;)&lt;br /&gt;
line_2_1000 = ax2.plot(results2_1000, &amp;quot;r&amp;quot;)&lt;br /&gt;
line_2_10000 = ax2.plot(results2_10000, &amp;quot;b&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
ax1.set_title(&amp;quot;Population = 1000&amp;quot;)&lt;br /&gt;
ax2.set_title(&amp;quot;Population = 10000&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
plt.title(&amp;quot;Mutation Meltdown&amp;quot;)&lt;br /&gt;
plt.show()&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Weigang&lt;br /&gt;
&amp;lt;syntaxhighlight lang=R&amp;quot;&amp;gt;&lt;br /&gt;
add.mutation &amp;lt;- function(genome, genome.length, mutation.rate) {&lt;br /&gt;
  mu.exp &amp;lt;- genome.length * mutation.rate;&lt;br /&gt;
  mu.num &amp;lt;- rpois(1, lambda = mu.exp);&lt;br /&gt;
  return(genome+mu.num);&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
ratchet &amp;lt;- function(pop.size=100, gamete.size=100, mu=1e-4, genome.length=1e4) {&lt;br /&gt;
  out &amp;lt;- data.frame(generation=numeric(), n0=numeric(), n1=numeric(), pop=numeric(), length=numeric());&lt;br /&gt;
  pop &amp;lt;- rep(0, pop.size) # initial all mutation-free&lt;br /&gt;
  g &amp;lt;- 1; # generation&lt;br /&gt;
  freq0 &amp;lt;- 1; # freq of zero-class&lt;br /&gt;
  while(freq0 &amp;gt; 0) {&lt;br /&gt;
    cat(&amp;quot;at generation&amp;quot;, g, &amp;quot;\n&amp;quot;);&lt;br /&gt;
    freq0 &amp;lt;- length(which(pop == 0));&lt;br /&gt;
    out &amp;lt;- rbind(out, data.frame(generation=g, n0=freq0/pop.size, pop=pop.size, length=genome.length));&lt;br /&gt;
    gametes &amp;lt;- sapply(1:gamete.size, function(x) {lapply(pop, function(x) add.mutation(x, genome.length, mu))});&lt;br /&gt;
    pop &amp;lt;- sample(gametes, pop.size);&lt;br /&gt;
    g &amp;lt;- g+1;&lt;br /&gt;
  }&lt;br /&gt;
  return(out);&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
out.long.genome.df &amp;lt;- ratchet(genome.length=1e4);&lt;br /&gt;
out.short.genome.df &amp;lt;- ratchet(genome.length=1e3);&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==March 31, 2017 &amp;quot;Drift-Mutation Balance&amp;quot; (4/14/2017)==&lt;br /&gt;
[[File:Drift-mutation.png|thumbnail]]&lt;br /&gt;
We will explore genetic drift of a DNA fragment with mutation under Wright-Fisher model. From last week&#039;s exercise, we conclude that a population will sooner or later lose genetic diversity (becoming fixed after ~2N generations), if no new alleles are generated (by e.g., mutation or migration). &lt;br /&gt;
&lt;br /&gt;
Mutation, in contrast, increases genetic diversity over time. Under neutrality (no natural selection against or for any mutation, e.g., on an intron sequence), the population will reach an equilibrium point when the loss of genetic diversity by drift is cancelled out by increase of genetic diversity by mutation.&lt;br /&gt;
&lt;br /&gt;
You job is to find this equilibrium point by simulation, given a population size (N) and a mutation rate (mu). The expected answer is pi=2N*mu, where pi is a measure of genetic diversity using DNA sequences, which is the average pairwise sequence differences within a population.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;font color=&amp;quot;blue&amp;quot;&amp;gt;Note: the following algorithm seems to be too complex to build at once. Also, too slow to run in R. Nonetheless, please try to write two R functions: (1)mutate.seq(seq, mut), with &amp;quot;seq&amp;quot; as a vector of bases and &amp;quot;mut&amp;quot; is the mutation rate, and (2) avg.seq.diff(pop), with &amp;quot;pop&amp;quot; as a list of DNA sequences&amp;lt;/font&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A suggested algorithm is:&lt;br /&gt;
# Start with a homogeneous population with N=100 identical DNA sequences (e.g., with a length of L=1e4 bases, or about 5 genes) [R hint: &amp;lt;code&amp;gt;dna &amp;lt;- sample(c(&amp;quot;a&amp;quot;, &amp;quot;t&amp;quot;, &amp;quot;c&amp;quot;, &amp;quot;g&amp;quot;), size=1e5, replace=T, prob = rep(0.25,4))&amp;lt;/code&amp;gt;]&lt;br /&gt;
# Write a mutation function, which will mutate the DNA based on Poisson process (since mutation is a rare event). For example, if mu=1e-4 per generation per base per individual (too high for real, but faster for results to converge), then each generation the expected number of mutations would be L * mu = 1 per individual per generation for our DNA segment. You would then simulate the random number of mutations by using the R function &amp;lt;code&amp;gt;num.mutations &amp;lt;- rpois(1,lambda=1)&amp;lt;/code&amp;gt;.&lt;br /&gt;
# Apply the mutation function for each individual DNA copy (a total of N=100) during gamete production (100 gametes for each individual) at each generation (for a total of G=1000 generations).&lt;br /&gt;
# Write another function to calculate, for each generation, instead of counting allele frequencies (as last week&#039;s problem), to calculate &amp;amp; output average pairwise differences among the 100 individuals.&lt;br /&gt;
# Finally, you would graph pi over generation.&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
| &lt;br /&gt;
* By John&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;python&amp;quot;&amp;gt;&lt;br /&gt;
import numpy as np&lt;br /&gt;
import matplotlib.pyplot as plt&lt;br /&gt;
from itertools import combinations&lt;br /&gt;
from random import sample&lt;br /&gt;
from numpy.random import choice&lt;br /&gt;
from numpy.random import poisson&lt;br /&gt;
&lt;br /&gt;
# global variables&lt;br /&gt;
nuc = np.array([&amp;quot;A&amp;quot;, &amp;quot;T&amp;quot;, &amp;quot;C&amp;quot;, &amp;quot;G&amp;quot;])&lt;br /&gt;
expected_mutation_rate = 0.00001 # there are 2 sequences per individual&lt;br /&gt;
seq_length = 10000&lt;br /&gt;
individual_pop = 100&lt;br /&gt;
gamete_rate = 100&lt;br /&gt;
generation = 1000&lt;br /&gt;
&lt;br /&gt;
def simulator(population, repro_rate, generation):&lt;br /&gt;
    observed_mutation_collection = []&lt;br /&gt;
    &lt;br /&gt;
    # Original Individuals&lt;br /&gt;
    original = choice(nuc, seq_length, 0.25)&lt;br /&gt;
    individual_total = [original for i in range(population)]&lt;br /&gt;
    &lt;br /&gt;
    for i in range(generation): # iterate over 1000 generations&lt;br /&gt;
        # Produce Gametes&lt;br /&gt;
        gametes = []&lt;br /&gt;
        for individual in individual_total:&lt;br /&gt;
            gametes += [individual for i in range(repro_rate)]&lt;br /&gt;
        gametes = np.array(gametes)&lt;br /&gt;
        gametes_pop = gametes.shape[0] # number of gametes: 100 * 100&lt;br /&gt;
        &lt;br /&gt;
        # Mutation&lt;br /&gt;
        mutation_number_arr = poisson(lam=seq_length * expected_mutation_rate, size=gametes_pop) # derive number of mutations per individual&lt;br /&gt;
        mutation_index_arr = [sample(range(seq_length), mutations) for mutations in mutation_number_arr] # get the index of mutation&lt;br /&gt;
        # Mutation: Replace with Mututated Base Pairs&lt;br /&gt;
        for i in range(gametes_pop): # iterate over all gametes&lt;br /&gt;
            for ix in mutation_index_arr[i]: # iterate over mutated base pair for each gamete&lt;br /&gt;
                gametes[i, ix] = choice(nuc[nuc != gametes[i, ix]], 1)[0] # locate mutation and alter the nuc with mutated one&lt;br /&gt;
        &lt;br /&gt;
        # Next generation of individuals&lt;br /&gt;
        individual_total = gametes[choice(range(10000), 100, replace=False)]&lt;br /&gt;
        &lt;br /&gt;
        # Calculate observed mutation rate&lt;br /&gt;
        num_combinations = 0&lt;br /&gt;
        total_diff_bases = 0&lt;br /&gt;
        for pair in combinations(range(len(individual_total)), 2): # aggregate mutations for all possible pairs of individuals&lt;br /&gt;
            total_diff_bases += len(np.where((individual_total[pair[0]] == individual_total[pair[1]]) == False)[0])&lt;br /&gt;
            num_combinations += 1&lt;br /&gt;
        observed_mutation_rate = float(total_diff_bases) / float(num_combinations) / seq_length&lt;br /&gt;
        observed_mutation_collection.append(observed_mutation_rate)&lt;br /&gt;
        &lt;br /&gt;
    return(observed_mutation_collection)&lt;br /&gt;
&lt;br /&gt;
# Simulation&lt;br /&gt;
result = simulator(100, 100, 1000)&lt;br /&gt;
&lt;br /&gt;
# Visualize results&lt;br /&gt;
plt.plot(result, &#039;b&#039;)&lt;br /&gt;
plt.xlabel(&amp;quot;Generation&amp;quot;)&lt;br /&gt;
plt.ylabel(&amp;quot;Mutation Rate&amp;quot;)&lt;br /&gt;
plt.title(&amp;quot;Drift-Mutation Balance&amp;quot;)&lt;br /&gt;
plt.show()&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt; &lt;br /&gt;
&lt;br /&gt;
*By Weigang&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
library(Biostrings) # a more compressed way to store and manipulate sequences&lt;br /&gt;
bases &amp;lt;- DNA_ALPHABET[1:4];&lt;br /&gt;
dna &amp;lt;- sample(bases, size = 1e5, replace = T);&lt;br /&gt;
library(ape) # to use the DNAbin methods&lt;br /&gt;
# function to mutate (Poisson process)&lt;br /&gt;
mutate.seq &amp;lt;- function(seq, mutation.rate) {&lt;br /&gt;
  mu.exp &amp;lt;- length(seq) * mutation.rate;&lt;br /&gt;
  mu.num &amp;lt;- rpois(1, lambda = mu.exp);&lt;br /&gt;
  if (mu.num &amp;gt; 0) {&lt;br /&gt;
    pos &amp;lt;- sample(1:length(seq), size = mu.num);&lt;br /&gt;
    for (j in 1:length(pos)) {&lt;br /&gt;
      current.base &amp;lt;- seq[pos[j]];&lt;br /&gt;
      mutated.base &amp;lt;-sample(bases[bases !=current.base)], size = 1);&lt;br /&gt;
      seq[pos[j]] &amp;lt;- mutated.base;&lt;br /&gt;
    }&lt;br /&gt;
  }&lt;br /&gt;
  return(seq)&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
# main routine&lt;br /&gt;
drift.mutation &amp;lt;- function(pop.size=100, generation.time=500, gamete.size=10, mu=1e-5) {&lt;br /&gt;
  out &amp;lt;- data.frame(generation=numeric(), pi=numeric()); # for storing outputs&lt;br /&gt;
  pop &amp;lt;- lapply(1:pop.size, function(x) dna)  # create the initial population&lt;br /&gt;
  for(i in 1:generation.time) { # for each generation&lt;br /&gt;
    cat(&amp;quot;at generation&amp;quot;, i, &amp;quot;\n&amp;quot;); # print progress&lt;br /&gt;
    gametes &amp;lt;- sapply(1:gamete.size, function(x) {lapply(pop, function(x) mutate.seq(x, mu))}); &lt;br /&gt;
    pop &amp;lt;- sample(gametes, pop.size); &lt;br /&gt;
    pop.bin &amp;lt;- as.DNAbin(pop); # change into DNAbin class to advantage of its dist.dna() function&lt;br /&gt;
    out &amp;lt;- rbind(out, data.frame(generation=i, pi=mean(dist.dna(pop.bin))));&lt;br /&gt;
  }&lt;br /&gt;
  out;&lt;br /&gt;
}&lt;br /&gt;
out.df &amp;lt;- drift.mutation(); # run function and save results into a data frame&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==March 18, 2017 &amp;quot;Genetic Drift&amp;quot; (Due 3/31/2017)==&lt;br /&gt;
[[File:Drift.png|thumbnail]]&lt;br /&gt;
[[File:Drift-founder.png|thumbnail]]&lt;br /&gt;
This is our first biological simulation. Mandatory assignment for all lab members (from interns to doctoral students). An expected result is shown in the graph.&lt;br /&gt;
&lt;br /&gt;
Task: Simulate the Wright-Fisher model of genetic drift as follows:&lt;br /&gt;
# Begin with an allele frequency p=0.5, and a pop of N=100 haploid individuals [Hint: &amp;lt;code&amp;gt;pop &amp;lt;- c(rep(&amp;quot;A&amp;quot;, 50), rep(&amp;quot;a&amp;quot;, 50))&amp;lt;/code&amp;gt;]&lt;br /&gt;
# Each individual produces 100 gametes, giving a total of 10,000 gametes [Hint: use a for loop with rep() function]&lt;br /&gt;
# Sample from the gamete pool another 100 to give rise to a new generation of individuals [Hint: use sample() function]&lt;br /&gt;
# Calculate allele frequency [Hint: use table() function]&lt;br /&gt;
# Repeat the above in succession for a total generation of g=1000 generations [Hint: create a function with three arguments, e.g., wright.fisher(pop.size, gamete.size, generation.time)]&lt;br /&gt;
# Plot allele frequency changes over generation time&lt;br /&gt;
# Be prepared to answer these questions:&lt;br /&gt;
## Why allele frequency fluctuate even without natural selection?&lt;br /&gt;
## What&#039;s the final fate of population, one allele left, or two alleles coexist indefinitely?&lt;br /&gt;
## Which population can maintain genetic polymorphism (with two alleles) longer?&lt;br /&gt;
## Which population gets fixed (only one allele remains) quicker?&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
|&lt;br /&gt;
* By Lili (May, 2019)&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
#Simple sampling, haploid, as a function &lt;br /&gt;
wright_fisher &amp;lt;- function(pop_size, gam_size, alle_frq, n_gen) {&lt;br /&gt;
  pop &amp;lt;- c(rep(&amp;quot;A&amp;quot;, pop_size*frq), rep(&amp;quot;a&amp;quot;, pop_size*frq))&lt;br /&gt;
  prob &amp;lt;- numeric(n_gen)&lt;br /&gt;
  for(time in 1:n_gen){&lt;br /&gt;
    gamt &amp;lt;- rep(pop, gam_size)&lt;br /&gt;
    pop &amp;lt;- sample(gamt, pop_size)&lt;br /&gt;
    prob[time] &amp;lt;- table(pop)[1]/pop_size&lt;br /&gt;
  }&lt;br /&gt;
  wf_df &amp;lt;- data.frame(generation=1:n_gen, probability=prob)&lt;br /&gt;
  plot(drift_df, type= &#039;l&#039;, main = &amp;quot;Wright-Fisher Process (Genetic Drift)&amp;quot;, xlab = &amp;quot;generation&amp;quot;, ylab = &amp;quot;allele frequency&amp;quot;)&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
wright_fisher(1000, 100, 0.5, 1000)&lt;br /&gt;
&lt;br /&gt;
#Version 2. Using binomial distribution with replicates, diploid&lt;br /&gt;
&lt;br /&gt;
pop_size &amp;lt;- c(50, 100, 1000, 5000)&lt;br /&gt;
alle_frq &amp;lt;- c(0.01, 0.1, 0.5, 0.8)&lt;br /&gt;
n_gen &amp;lt;- 100&lt;br /&gt;
n_reps &amp;lt;- 50&lt;br /&gt;
genetic_drift &amp;lt;- data.frame()&lt;br /&gt;
&lt;br /&gt;
for(N in pop_size){&lt;br /&gt;
  for(p in alle_frq){&lt;br /&gt;
    p0 &amp;lt;- p&lt;br /&gt;
    for(j in 1:n_gen){&lt;br /&gt;
      X &amp;lt;- rbinom(n_reps, 2*N, p)&lt;br /&gt;
      p &amp;lt;- X/(2*N)&lt;br /&gt;
      rows &amp;lt;- data.frame(replicate= 1:n_reps, pop=rep(N, n_reps), gen=rep(j, n_reps), frq=rep(p0, n_reps), prob=p )&lt;br /&gt;
      genetic_drift &amp;lt;- rbind(genetic_drift, rows)&lt;br /&gt;
    }&lt;br /&gt;
  }&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
library(ggplot2)&lt;br /&gt;
ggplot(genetic_drift, aes(x=gen, y=prob, group=replicate)) + geom_path(alpha= .5) + facet_grid(pop ~ frq) + guides(colour=FALSE)&lt;br /&gt;
&lt;br /&gt;
# 3rd version: trace ancestry (instead of allele frequency)&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By John&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;python&amp;quot;&amp;gt;&lt;br /&gt;
#Python&lt;br /&gt;
import numpy as np&lt;br /&gt;
import sys&lt;br /&gt;
from random import sample&lt;br /&gt;
import matplotlib.pyplot as plt&lt;br /&gt;
 &lt;br /&gt;
def simulator(gametes_rate, next_individuals, generations):&lt;br /&gt;
    individuals = [1 for i in range(50)] + [0 for j in range(50)]&lt;br /&gt;
    frequency = []&lt;br /&gt;
    for generation in range(generations):&lt;br /&gt;
        gametes = []&lt;br /&gt;
        for individual in individuals:&lt;br /&gt;
            gametes += [individual for i in range(gametes_rate)]&lt;br /&gt;
        individuals = sample(gametes, next_individuals)&lt;br /&gt;
        frequency.append(np.count_nonzero(individuals) / len(individuals))&lt;br /&gt;
    return(frequency)&lt;br /&gt;
 &lt;br /&gt;
N_100 = simulator(100, 100, 1000)&lt;br /&gt;
N_1000 = simulator(100, 1000, 1000)&lt;br /&gt;
 &lt;br /&gt;
# Create plots with pre-defined labels.&lt;br /&gt;
# Alternatively,pass labels explicitly when calling `legend`.&lt;br /&gt;
fig, ax = plt.subplots()&lt;br /&gt;
ax.plot(N_100, &#039;r&#039;, label=&#039;N=100&#039;)&lt;br /&gt;
ax.plot(N_1000, &#039;b&#039;, label=&#039;N=1000&#039;)&lt;br /&gt;
 &lt;br /&gt;
# Add x, y labels and title&lt;br /&gt;
plt.ylim(-0.1, 1.1)&lt;br /&gt;
plt.xlabel(&amp;quot;Generation&amp;quot;)&lt;br /&gt;
plt.ylabel(&amp;quot;Frequency&amp;quot;)&lt;br /&gt;
plt.title(&amp;quot;Wright-Fisher Model&amp;quot;)&lt;br /&gt;
 &lt;br /&gt;
# Now add the legend with some customizations.&lt;br /&gt;
legend = ax.legend(loc=&#039;upper right&#039;, shadow=True)&lt;br /&gt;
 &lt;br /&gt;
# The frame is matplotlib.patches.Rectangle instance surrounding the legend.&lt;br /&gt;
frame = legend.get_frame()&lt;br /&gt;
frame.set_facecolor(&#039;0.90&#039;)&lt;br /&gt;
 &lt;br /&gt;
# Set the fontsize&lt;br /&gt;
for label in legend.get_texts():&lt;br /&gt;
    label.set_fontsize(&#039;large&#039;)&lt;br /&gt;
 &lt;br /&gt;
for label in legend.get_lines():&lt;br /&gt;
    label.set_linewidth(1.5)  # the legend line width&lt;br /&gt;
plt.show()&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;scala&amp;quot;&amp;gt;&lt;br /&gt;
#Scala&lt;br /&gt;
import scala.util.Random&lt;br /&gt;
import scala.collection.mutable.ListBuffer&lt;br /&gt;
 &lt;br /&gt;
val gamete_rate = 100&lt;br /&gt;
val offspr_rate = 100&lt;br /&gt;
val generations = 1000&lt;br /&gt;
 &lt;br /&gt;
var frequency: List[Double] = List()&lt;br /&gt;
 &lt;br /&gt;
var individuals = List.fill(50)(0) ++ List.fill(50)(1)&lt;br /&gt;
 &lt;br /&gt;
for(generation &amp;lt;- 1 to generations){&lt;br /&gt;
  var gametes = individuals.map(x =&amp;gt; List.fill(gamete_rate)(x)).flatten&lt;br /&gt;
  individuals = Random.shuffle(gametes).take(offspr_rate)&lt;br /&gt;
  frequency = frequency :+ individuals.count(_ == 1).toDouble / offspr_rate&lt;br /&gt;
}&lt;br /&gt;
 &lt;br /&gt;
print(frequency)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;java&amp;quot;&amp;gt;&lt;br /&gt;
#Spark&lt;br /&gt;
val gamete_rate = 100&lt;br /&gt;
val offspr_rate = 100&lt;br /&gt;
val generations = 300&lt;br /&gt;
 &lt;br /&gt;
var frequency: List[Double] = List()&lt;br /&gt;
var individuals = sc.parallelize(List.fill(50)(&amp;quot;0&amp;quot;) ++ List.fill(50)(&amp;quot;1&amp;quot;))&lt;br /&gt;
 &lt;br /&gt;
for(generation &amp;lt;- 1 to generations){&lt;br /&gt;
  val gametes = individuals.flatMap(x =&amp;gt; (x * gamete_rate).split(&amp;quot;&amp;quot;).tail)&lt;br /&gt;
  individuals = sc.parallelize(gametes.takeSample(false, offspr_rate))&lt;br /&gt;
  val count = individuals.countByValue&lt;br /&gt;
  frequency = frequency :+ count(&amp;quot;1&amp;quot;).toDouble / offspr_rate&lt;br /&gt;
}&lt;br /&gt;
 &lt;br /&gt;
print(frequency)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Brian&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
Wright&amp;lt;-function(pop.size,gam.size,generation) {&lt;br /&gt;
 if( pop.size%%2==0) {&lt;br /&gt;
 &lt;br /&gt;
  pop&amp;lt;-c(rep(&amp;quot;A&amp;quot;,pop.size*.5),rep(&amp;quot;a&amp;quot;,pop.size))&lt;br /&gt;
  frequency&amp;lt;-.5&lt;br /&gt;
  time&amp;lt;-0&lt;br /&gt;
  geneticDrift&amp;lt;-data.frame(frequency,time)&lt;br /&gt;
  for (i in 1:generation) {&lt;br /&gt;
    largePop&amp;lt;-rep(pop,gam.size)&lt;br /&gt;
    samplePop&amp;lt;-sample(largePop,100)&lt;br /&gt;
    cases&amp;lt;-table(samplePop)&lt;br /&gt;
    if (cases[1]&amp;lt;100 &amp;amp;&amp;amp; cases[2]&amp;lt;100 ) &lt;br /&gt;
        {propA&amp;lt;-(cases[1]/100)&lt;br /&gt;
        geneticDrift&amp;lt;-rbind(geneticDrift,c(propA,i))&lt;br /&gt;
        pop&amp;lt;-c(rep(&amp;quot;A&amp;quot;,cases[2]),rep(&amp;quot;a&amp;quot;,cases[1]))&lt;br /&gt;
  }  &lt;br /&gt;
}&lt;br /&gt;
  plot(geneticDrift$frequency~geneticDrift$time,type=&amp;quot;b&amp;quot;, main=&amp;quot;Genetic Drift N=1000&amp;quot;, xlab=&amp;quot;time&amp;quot;,ylab=&amp;quot;Proportion Pop A&amp;quot;,col=&amp;quot;red&amp;quot;, pch=18 )&lt;br /&gt;
}&lt;br /&gt;
 if(pop.size%%2==1 ) {&lt;br /&gt;
   print(&amp;quot;Initial Population should be even. Try again&amp;quot;)&lt;br /&gt;
 }&lt;br /&gt;
}&lt;br /&gt;
Wright(1000,100,1000)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Jamila&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
Genetic_code = function(t,R){ &lt;br /&gt;
  N&amp;lt;- 100&lt;br /&gt;
  p&amp;lt;- 0.5&lt;br /&gt;
  frequency&amp;lt;-as.numeric();&lt;br /&gt;
  for (i in 1:t){&lt;br /&gt;
    A1=rbiom(1,2*N,p)&lt;br /&gt;
    p=A1/(N*2); &lt;br /&gt;
    frequency[length(frequency)+1]&amp;lt;-p; &lt;br /&gt;
  }&lt;br /&gt;
  plot(frequency, type=&amp;quot;1&amp;quot;,ylim=c(0,1),col=3,xlab=&amp;quot;t&amp;quot;,ylab=expression(p(A[1])))&lt;br /&gt;
  &lt;br /&gt;
  }&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Sharon&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
p=0.5&lt;br /&gt;
N=100&lt;br /&gt;
g=0&lt;br /&gt;
pop &amp;lt;- c(rep(&amp;quot;A&amp;quot;, 50), rep(&amp;quot;a&amp;quot;, 50)) #create a population of 2 alleles&lt;br /&gt;
data&amp;lt;-data.frame(g,p) #a set of variables of the same # of rows &lt;br /&gt;
for (i in 1:1000) { #generation &lt;br /&gt;
  gam_pl &amp;lt;-rep(pop, 100) #gamete pool&lt;br /&gt;
  gam_sam &amp;lt;-sample(gam_pl, 100)  #sample from gamete pool&lt;br /&gt;
  tab_all&amp;lt;-table(gam_sam) #table ps sample&lt;br /&gt;
  all_freq&amp;lt;-tab_all[1]/100 #get allele frequency&lt;br /&gt;
  data&amp;lt;-rbind(data, c(i,all_freq))&lt;br /&gt;
  pop&amp;lt;-c(rep(&amp;quot;A&amp;quot;,tab_all[1]))&lt;br /&gt;
}&lt;br /&gt;
plot(data$g,data$p)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Sipa&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
pop &amp;lt;- c(rep(&amp;quot;A&amp;quot;, 50), rep(&amp;quot;a&amp;quot;, 50))&lt;br /&gt;
wright_fisher &amp;lt;- function(N,gen_time,gam) {&lt;br /&gt;
  N=N&lt;br /&gt;
  gen_time = gen_time&lt;br /&gt;
  x = numeric(gen_time) &lt;br /&gt;
  x[1] = gam&lt;br /&gt;
  for (i in 2:1000) {&lt;br /&gt;
  k=(x[i-1])/N&lt;br /&gt;
  n=seq(0,N,1)&lt;br /&gt;
  prob=dbinom(n,N,k)&lt;br /&gt;
  x[i]=sample(0:N, 1, prob=prob)&lt;br /&gt;
  &lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
plot(x[1:gen_time], type=&amp;quot;l&amp;quot;, pch=10)&lt;br /&gt;
&lt;br /&gt;
}&lt;br /&gt;
pool &amp;lt;- rep(pop, times = 100)&lt;br /&gt;
s_pool &amp;lt;- sample(pool, size = 100, replace = F)&lt;br /&gt;
table(s_pool)&lt;br /&gt;
wright_fisher2 &amp;lt;- function(pop_size,al_freq,gen_time) {&lt;br /&gt;
  pop &amp;lt;- c(rep(&amp;quot;A&amp;quot;, pop_size/2), rep(&amp;quot;a&amp;quot;, pop_size/2))&lt;br /&gt;
  pool &amp;lt;-rep(pop, times=100)&lt;br /&gt;
  s_pool &amp;lt;- sample(pool, size = 100, replace = T)&lt;br /&gt;
  for(i in 1:gen_time)&lt;br /&gt;
    s_pool &amp;lt;- sample(sample(pool, size = pop_size, replace = F))&lt;br /&gt;
  a.f &amp;lt;- table(s_pool)[1]/100&lt;br /&gt;
  return(a.f)&lt;br /&gt;
  }&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Nicolette&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
N=100 &lt;br /&gt;
p=0.5 &lt;br /&gt;
g=1000 &lt;br /&gt;
sim=100 &lt;br /&gt;
Genetic_drift=array(0, dim=c(g,sim))&lt;br /&gt;
Genetic_drift[1,]=rep(N*p,sim)&lt;br /&gt;
for(i in 1:sim) {&lt;br /&gt;
  for(j in 2:g){&lt;br /&gt;
    X[j,i]=rbinom(1,N,prob=X[j-1,i]/N)&lt;br /&gt;
  }&lt;br /&gt;
}&lt;br /&gt;
Genetic_drift=data.frame(X/N)&lt;br /&gt;
matplot(1:1000, (X/N), type=&amp;quot;l&amp;quot;,ylab=&amp;quot;allele_frequency&amp;quot;,xlab=&amp;quot;generations&amp;quot;)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* By Weigang&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
# output frequency:&lt;br /&gt;
wright.fisher &amp;lt;- function(pop.size, generation.time=1000, gametes.per.ind=100) {&lt;br /&gt;
  out.df &amp;lt;- data.frame(gen=numeric(), freq=numeric(), pop.size=numeric());&lt;br /&gt;
  pop &amp;lt;- c(rep(&amp;quot;A&amp;quot;, pop.size/2), rep(&amp;quot;a&amp;quot;, pop.size/2)); # initial pop&lt;br /&gt;
  for(i in 1:generation.time) {&lt;br /&gt;
    gametes &amp;lt;- unlist(lapply(pop, function(x) rep(x, gametes.per.ind)));&lt;br /&gt;
    pop &amp;lt;- sample(gametes, pop.size); # sample gamete pool without replacement&lt;br /&gt;
    freq.a &amp;lt;- table(pop)[1]/pop.size; # frequency of &amp;quot;a&amp;quot;&lt;br /&gt;
    out.df &amp;lt;- rbind(out.df, c(gen=i, freq=freq.a, pop=pop.size, rep=rep));&lt;br /&gt;
  }&lt;br /&gt;
  out.list;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
out.1e2&amp;lt;-wright.fisher(100)&lt;br /&gt;
out.1e3&amp;lt;-wright.fisher(1000)&lt;br /&gt;
# Make Figure 1 above&lt;br /&gt;
plot(out.1e2[,1], out.1e2[,2], type = &amp;quot;l&amp;quot;, las=1, ylim=c(0,1), xlab=&amp;quot;generation&amp;quot;, ylab=&amp;quot;allele frequency&amp;quot;, main=&amp;quot;Wright-Fisher Process (Genetic Drift)&amp;quot;, col=3, lwd=2)&lt;br /&gt;
lines(out.1e3[,1], out.1e3[,2], type = &amp;quot;l&amp;quot;, col=2, lwd=2)&lt;br /&gt;
legend(400, 0.2, c(&amp;quot;N=100&amp;quot;, &amp;quot;N=1,000&amp;quot;), lty=1, col=3:2, cex=0.75)&lt;br /&gt;
abline(h=0.5, col=&amp;quot;gray&amp;quot;, lty=2)&lt;br /&gt;
&lt;br /&gt;
# Second function to track founders (not frequency)&lt;br /&gt;
wright.fisher.2 &amp;lt;- function(pop.size, generation.time=1000, gametes.per.ind=100) {&lt;br /&gt;
  pop &amp;lt;- 1:100 # label founders&lt;br /&gt;
  out.list &amp;lt;- list(pop);&lt;br /&gt;
  for(i in 1:generation.time) {&lt;br /&gt;
    gametes &amp;lt;- unlist(lapply(pop, function(x) rep(x, gametes.per.ind)));&lt;br /&gt;
    pop &amp;lt;- sample(gametes, pop.size); # sample gamete pool without replacement&lt;br /&gt;
    out.list[[i+1]] &amp;lt;- pop;&lt;br /&gt;
  }&lt;br /&gt;
  out.list;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
out.1ist&amp;lt;-wright.fisher.2(100)&lt;br /&gt;
# plot loss of genetic diversity by tracking surviving founder lineages&lt;br /&gt;
plot(x=1:100, y=rep(1, 100), ylim=c(0,100), xlim=c(1,600), type=&amp;quot;n&amp;quot;, las=1, main=&amp;quot;Drift: Suvival of founder lineages\n(N=100)&amp;quot;, cex.main=0.75, xlab=&amp;quot;generation&amp;quot;, ylab=&amp;quot;founders&amp;quot;)&lt;br /&gt;
for (i in 1:600) { points(rep(i,100), 1:100, col=colors()[out.list[[i]]], cex=0.5) }&lt;br /&gt;
for (i in 1:600) {points(rep(i,100), out.list[[i]], cex=0.2, col=&amp;quot;yellow&amp;quot;) }&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== March 11, 2017 &amp;quot;Stoplights&amp;quot; (Due 3/18/2017)==&lt;br /&gt;
[[File:Stoplight-1.png|framed|right]]&lt;br /&gt;
* Source: Paul Nahin (2008). &amp;quot;Digital Dice&amp;quot;, Problem 18.&lt;br /&gt;
* Challenge: How many red lights, on average, will you have to wait for on your journey from a city block m streets and n avenues away from Belfer [with coordinates (1,1)]? (assuming equal probability for red and green lights)&lt;br /&gt;
* Note that one has only wait for green light when walking along either the north side of 69 Street or east side of 1st Avenue. On all other intersections, one can walk non-stop without waiting for green light by crossing in the other direction if a red light is on.&lt;br /&gt;
* Formulate your solution with the following steps:&lt;br /&gt;
# Start from (m+1,n+1) corner and end at (1,1) corner&lt;br /&gt;
# The average number of red lights for m=n=0 is zero&lt;br /&gt;
# Find the average number of red lights for m=n=1 by simulating the walk 100 times&lt;br /&gt;
# Increment m &amp;amp; n by 1 (but keep m=n), until m=n=1000&lt;br /&gt;
# Plot average number of red lights by m (or n).&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
| &lt;br /&gt;
*By John&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;python&amp;quot;&amp;gt;&lt;br /&gt;
import numpy as np&lt;br /&gt;
import matplotlib.pyplot as plt&lt;br /&gt;
import pandas as pd&lt;br /&gt;
from pandas import DataFrame&lt;br /&gt;
from random import sample&lt;br /&gt;
from itertools import combinations, permutations&lt;br /&gt;
from numpy import count_nonzero, array&lt;br /&gt;
 &lt;br /&gt;
def red_light(av, st):&lt;br /&gt;
    traffic_light = [&amp;quot;red_st&amp;quot;, &amp;quot;red_av&amp;quot;]&lt;br /&gt;
    total_cross = av + st - 1&lt;br /&gt;
    while av != 0 and st != 0:&lt;br /&gt;
        if sample(traffic_light, 1)[0] == &amp;quot;red_st&amp;quot;:&lt;br /&gt;
            av -= 1&lt;br /&gt;
        else:&lt;br /&gt;
            st -= 1&lt;br /&gt;
    rest_cross = av if av != 0 else st&lt;br /&gt;
    return(count_nonzero([sample([0, 1], 1) for corss in range(rest_cross)]))&lt;br /&gt;
 &lt;br /&gt;
df = DataFrame(0, index=np.arange(10), columns=np.arange(10))&lt;br /&gt;
 &lt;br /&gt;
simulation = 1000&lt;br /&gt;
for av in df.index:&lt;br /&gt;
    for st in df.index:&lt;br /&gt;
        df.loc[av, st] = sum(array([[red_light(av, st)] for n in range(simulation)])) / simulation&lt;br /&gt;
        &lt;br /&gt;
plt.imshow(df, cmap=&#039;hot&#039;, interpolation=&#039;nearest&#039;)&lt;br /&gt;
plt.xticks(np.arange(1, 10))&lt;br /&gt;
plt.yticks(np.arange(1, 10))&lt;br /&gt;
plt.title(&amp;quot;Average Numbers Waiting for Red Light&amp;quot;)&lt;br /&gt;
plt.xlabel(&amp;quot;Number of Av&amp;quot;)&lt;br /&gt;
plt.ylabel(&amp;quot;Number of St&amp;quot;)&lt;br /&gt;
plt.show()&lt;br /&gt;
 &lt;br /&gt;
df&lt;br /&gt;
 &lt;br /&gt;
# Recursive Function for Each Trial&lt;br /&gt;
&lt;br /&gt;
from sys import stdout&lt;br /&gt;
av = 30&lt;br /&gt;
st = 30&lt;br /&gt;
n_space = [0]&lt;br /&gt;
traffic_light = [&amp;quot;red_st&amp;quot;, &amp;quot;red_av&amp;quot;]&lt;br /&gt;
def route(av, st):&lt;br /&gt;
    if not av == st == 0:&lt;br /&gt;
        if sample(traffic_light, 1)[0] == &amp;quot;red_st&amp;quot;:&lt;br /&gt;
            if av == 0:&lt;br /&gt;
                stdout.write(&amp;quot;w_&amp;quot;)&lt;br /&gt;
                n_space[0] += 1&lt;br /&gt;
                route(av, st - 1)&lt;br /&gt;
            else:&lt;br /&gt;
                stdout.write(&amp;quot;\n&amp;quot; + &amp;quot; &amp;quot; * n_space[0] + &amp;quot;|&amp;quot;)&lt;br /&gt;
                route(av - 1, st)&lt;br /&gt;
        else:&lt;br /&gt;
            if st == 0:&lt;br /&gt;
                stdout.write(&amp;quot;\n&amp;quot; + &amp;quot; &amp;quot; * n_space[0] + &amp;quot;w&amp;quot; + &amp;quot;\n&amp;quot; + &amp;quot; &amp;quot; * n_space[0] + &amp;quot;|&amp;quot;)&lt;br /&gt;
                route(av - 1, st)&lt;br /&gt;
            else:&lt;br /&gt;
                stdout.write(&amp;quot;_&amp;quot;)&lt;br /&gt;
                n_space[0] += 1&lt;br /&gt;
                route(av, st - 1)&lt;br /&gt;
route(av, st)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Weigang&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;perl&amp;quot;&amp;gt;&lt;br /&gt;
#!/usr/bin/env perl&lt;br /&gt;
use strict;&lt;br /&gt;
use warnings;&lt;br /&gt;
&lt;br /&gt;
# Use recursion&lt;br /&gt;
my ($distx, $disty) = @ARGV;&lt;br /&gt;
my $num_red = 0; # keep red-light counts&lt;br /&gt;
print &amp;quot;-&amp;quot; x $distx, &amp;quot;\n&amp;quot;; # print a starting line&lt;br /&gt;
&amp;amp;walk($distx, $disty, \$num_red); # pass reference not value&lt;br /&gt;
exit;&lt;br /&gt;
&lt;br /&gt;
sub walk {&lt;br /&gt;
    my ($x, $y, $ref) = @_;&lt;br /&gt;
    my $ct = $$ref;&lt;br /&gt;
    my $prob_red;&lt;br /&gt;
    if ($x == 1 &amp;amp;&amp;amp; $y == 1) { # Reached destination&lt;br /&gt;
	print &amp;quot;*\n&amp;quot;;&lt;br /&gt;
	print &amp;quot;-&amp;quot; x $distx, &amp;quot;\n&amp;quot;; # print the ending line&lt;br /&gt;
	print &amp;quot;reached Belfer after waiting for &amp;quot;, $ct, &amp;quot; red lights\n&amp;quot;;&lt;br /&gt;
    }&lt;br /&gt;
&lt;br /&gt;
    if ($x == 1 &amp;amp;&amp;amp; $y &amp;gt; 1) { # Reached right-side end&lt;br /&gt;
	$prob_red = rand();&lt;br /&gt;
	if ($prob_red &amp;lt; 0.5) { # red light, wait&lt;br /&gt;
	    $ct++;&lt;br /&gt;
	    print &amp;quot;w\n&amp;quot;, &amp;quot; &amp;quot; x ($distx-1);&lt;br /&gt;
	} else {&lt;br /&gt;
	    print &amp;quot;|\n&amp;quot;, &amp;quot; &amp;quot; x ($distx-1);&lt;br /&gt;
	}&lt;br /&gt;
	&amp;amp;walk($x, $y-1, \$ct);&lt;br /&gt;
    }&lt;br /&gt;
&lt;br /&gt;
    if ($x &amp;gt; 1 &amp;amp;&amp;amp; $y == 1) { # Reached bottom&lt;br /&gt;
	$prob_red = rand();&lt;br /&gt;
	if ($prob_red &amp;lt; 0.5) { # red light, wait&lt;br /&gt;
	    $ct++;&lt;br /&gt;
	    print &amp;quot;w&amp;quot;;&lt;br /&gt;
	} else { print &amp;quot;-&amp;quot;}&lt;br /&gt;
	&amp;amp;walk($x-1, $y, \$ct);&lt;br /&gt;
    }&lt;br /&gt;
&lt;br /&gt;
    if ($x &amp;gt; 1 &amp;amp;&amp;amp; $y &amp;gt; 1) {&lt;br /&gt;
	my $prob_across = rand(); # prob of walking right with green light&lt;br /&gt;
	if ($prob_across &amp;gt;= 0.5) { # move one block right&lt;br /&gt;
	    print &amp;quot;-&amp;quot;;&lt;br /&gt;
	    &amp;amp;walk($x-1, $y, \$ct);&lt;br /&gt;
	} else { # red light, move one block down&lt;br /&gt;
	    print &amp;quot;|\n&amp;quot;, &amp;quot; &amp;quot; x ($distx-$x);&lt;br /&gt;
	    &amp;amp;walk($x, $y-1, \$ct);&lt;br /&gt;
	}&lt;br /&gt;
    }&lt;br /&gt;
}&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Jeff&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;python&amp;quot;&amp;gt;&lt;br /&gt;
# In[26]:&lt;br /&gt;
def walk (walker):&lt;br /&gt;
    import random&lt;br /&gt;
    &lt;br /&gt;
    def just_walk (x):                                  #[1]&lt;br /&gt;
        if random.choice([&amp;quot;red_m_direction&amp;quot;, &amp;quot;green_m&amp;quot;])==&amp;quot;green_m&amp;quot;: x[&amp;quot;loca&amp;quot;][&amp;quot;m&amp;quot;] -=1&lt;br /&gt;
        else: x[&amp;quot;loca&amp;quot;][&amp;quot;n&amp;quot;] -=1&lt;br /&gt;
        return x    &lt;br /&gt;
    def may_have_to_wait (dist,waited):                 #[2]&lt;br /&gt;
        if random.choice([&amp;quot;green&amp;quot;,&amp;quot;red&amp;quot;])==&amp;quot;red&amp;quot;: waited += 1&lt;br /&gt;
        else: dist -= 1&lt;br /&gt;
        return (dist,waited)    &lt;br /&gt;
&lt;br /&gt;
    while (0 not in walker[&amp;quot;loca&amp;quot;].values()):         #start walking&lt;br /&gt;
        walker = just_walk(walker) &lt;br /&gt;
    while (walker[&amp;quot;loca&amp;quot;][&amp;quot;m&amp;quot;] !=0):       # if n =0 and m != 0&lt;br /&gt;
        (walker[&amp;quot;loca&amp;quot;][&amp;quot;m&amp;quot;], walker[&amp;quot;waited&amp;quot;]) = may_have_to_wait(walker[&amp;quot;loca&amp;quot;][&amp;quot;m&amp;quot;],walker[&amp;quot;waited&amp;quot;]) &lt;br /&gt;
    while (walker[&amp;quot;loca&amp;quot;][&amp;quot;n&amp;quot;] !=0):       # if m =0 and n != 0&lt;br /&gt;
        (walker[&amp;quot;loca&amp;quot;][&amp;quot;n&amp;quot;], walker[&amp;quot;waited&amp;quot;]) = may_have_to_wait(walker[&amp;quot;loca&amp;quot;][&amp;quot;n&amp;quot;],walker[&amp;quot;waited&amp;quot;]) &lt;br /&gt;
&lt;br /&gt;
    return walker[&amp;quot;waited&amp;quot;]&lt;br /&gt;
&lt;br /&gt;
def given_distance(m,n,rep): &lt;br /&gt;
    waited = list()&lt;br /&gt;
    for x in range(rep):&lt;br /&gt;
        walker={&amp;quot;loca&amp;quot;:{&amp;quot;m&amp;quot;:m,&amp;quot;n&amp;quot;:n}, &amp;quot;waited&amp;quot;:0}     #walker variable&lt;br /&gt;
        waited.append(walk(walker))      &lt;br /&gt;
    return sum(waited)/len(waited)&lt;br /&gt;
# main&lt;br /&gt;
result=list()&lt;br /&gt;
for d in range(1000):                                  # set walking distance here&lt;br /&gt;
        result.append(given_distance (d,d,100))      # set rep here&lt;br /&gt;
import matplotlib.pyplot as plt                     # ploting &lt;br /&gt;
plt.plot(result); plt.show()  &lt;br /&gt;
&lt;br /&gt;
#[1] no reason to wait for ANY red before one of the distances (m or n) is exhausted. Assuming when light is green for m, it must be red for n, vise versa. &lt;br /&gt;
#[2] when one of the directions (m or n = 0), waiting cannot be avoided, start counting. &lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== March 3, 2017 &amp;quot;PiE&amp;quot; (Due 3/10/2017)==&lt;br /&gt;
* Source: Paul Nahin (2008). &amp;quot;Dueling Idiots&amp;quot;, Problem 5.&lt;br /&gt;
* Challenge: obtain numerical values of Pi and E by simulations&lt;br /&gt;
* Simulate pi by [[File:Pi-sim.png|thumbnail]]&lt;br /&gt;
# Randomly generate 10,000 pairs of uniformly distributed numbers from 0 and 1 (simulating throwing darts onto the unit square shown at right)&lt;br /&gt;
# Count the number of points enclosed within the quarter-circle&lt;br /&gt;
# Calculate the pi value from this proportion&lt;br /&gt;
* Simulate e by&lt;br /&gt;
# Generate N random numbers from 0 to 1&lt;br /&gt;
# Divide into N equal-width bins between 0 and 1&lt;br /&gt;
# Count the number of bins Z that receive none of the random numbers&lt;br /&gt;
# Obtain e ~ N/Z (based on binomial sampling formula)&lt;br /&gt;
# Simulate N=1e2, 1e3, and 1e4&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
| &lt;br /&gt;
* By Weigang&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
# simulate pi:&lt;br /&gt;
darts &amp;lt;- sapply(1:1000, function(x) { coords&amp;lt;- runif(2); return(ifelse(coords[1]^2+coords[2]^2 &amp;lt;= 1, 1,0)) })&lt;br /&gt;
pi &amp;lt;- 4*mean(darts)&lt;br /&gt;
&lt;br /&gt;
# simulate e&lt;br /&gt;
N &amp;lt;- 1e3;&lt;br /&gt;
n &amp;lt;- runif(N);&lt;br /&gt;
cts &amp;lt;- numeric();&lt;br /&gt;
for(i in 1:(N-1)) {&lt;br /&gt;
  left &amp;lt;- i/N;&lt;br /&gt;
  right &amp;lt;- (i+1)/N;&lt;br /&gt;
  cts[i] &amp;lt;- length(which(n&amp;gt;=left &amp;amp; n &amp;lt; right))&lt;br /&gt;
}&lt;br /&gt;
e &amp;lt;- N/length(which(cts==0))&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Nicolette&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
#Simulate pi&lt;br /&gt;
random &amp;lt;- 0; square &amp;lt;- 0;&lt;br /&gt;
for (i in 1:1000){&lt;br /&gt;
  random[[i]]&amp;lt;- runif(1,0,1)&lt;br /&gt;
square[[i]]&amp;lt;- sqrt(1-(random[i])^2)&lt;br /&gt;
}&lt;br /&gt;
plot(random, square)&lt;br /&gt;
areaofqc &amp;lt;- (pi/4)&lt;br /&gt;
ranleqc &amp;lt;- length(which(random &amp;lt;=areaofqc))&lt;br /&gt;
squleqc &amp;lt;- length(which(square&amp;lt;=areaofqc))&lt;br /&gt;
&lt;br /&gt;
#Simulate e&lt;br /&gt;
e &amp;lt;- 0;&lt;br /&gt;
for(N in 1:1000) {&lt;br /&gt;
  numbers&amp;lt;-runif(N)&lt;br /&gt;
bin.size&amp;lt;-1/N&lt;br /&gt;
  non.empty&amp;lt;-as.integer(numbers/bin.size)&lt;br /&gt;
  z.empt&amp;lt;- N - length(table(non.empty))&lt;br /&gt;
  e&amp;lt;-c(e, N/z.empt)&lt;br /&gt;
  }&lt;br /&gt;
plot(e)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Brian&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
#When you throw a dart at the a unit square dartboard the Probability of hitting the portion of the circle radius with center=(0,0) inside the unit square is  pi/4.  We can say that hitting the 1/4 circle within the unit square with a dart is a bernoulli random variable where p=pi/4. Further, we can say E(bernoulli=hit)=pi/4. &lt;br /&gt;
# Imagine you are scoring the game of darts as follows- 1 point if you throw in the 1/4 circle and 0 points if you miss this space. &lt;br /&gt;
# If you throw 10000 darts you just did 10000 iid bernoulli trials. By the law of large numbers if we count the number of hits to the 1/4 circle of radius 1 and divide by number of darts thrown we will get something pretty close to E(beroulli) . Multiply that number by 4 and you have an estimate of Pi.  &lt;br /&gt;
&lt;br /&gt;
## Generate 10,000 pairs of points in the unit square with uniform distribution&lt;br /&gt;
&lt;br /&gt;
x&amp;lt;-c(runif(1000000, min = 0, max = 1))&lt;br /&gt;
y&amp;lt;-c(runif(1000000, min = 0, max = 1))&lt;br /&gt;
point&amp;lt;-data.frame(x,y)&lt;br /&gt;
&lt;br /&gt;
plot(point$x,point$y)&lt;br /&gt;
point.sub&amp;lt;-subset(point,y&amp;lt;=sqrt(1-x^2))&lt;br /&gt;
&lt;br /&gt;
plot(point.sub$x,point.sub$y)&lt;br /&gt;
&lt;br /&gt;
z&amp;lt;-4*nrow(point.sub)/1000000&lt;br /&gt;
z&lt;br /&gt;
error&amp;lt;-pi-z&lt;br /&gt;
error&lt;br /&gt;
&lt;br /&gt;
## simulating exp&lt;br /&gt;
#It&#039;s a well that a binomial with large n and tiny p is a good estimate of the poisson random For this estimate lambda=np.  In this case n=10,000 and p=the probability of falling into a particular unit=1/10000....... &lt;br /&gt;
 simexp&amp;lt;-function (n) {&lt;br /&gt;
a&amp;lt;-c(runif(n,min=0,max=1))&lt;br /&gt;
b&amp;lt;-c()&lt;br /&gt;
&lt;br /&gt;
for (i in 1:n) {&lt;br /&gt;
  c&amp;lt;-subset(a,(i-1)/n&amp;lt;a &amp;amp; a&amp;lt;=i/n)&lt;br /&gt;
  d&amp;lt;-length(c)&lt;br /&gt;
  b&amp;lt;-append(b,d)&lt;br /&gt;
}&lt;br /&gt;
bzero&amp;lt;-subset(b,b==0)&lt;br /&gt;
length(bzero)&lt;br /&gt;
print(n/length(bzero))&lt;br /&gt;
}&lt;br /&gt;
simexp(1000)&lt;br /&gt;
simexp(10000)&lt;br /&gt;
simexp(100000)&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By John&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
#python code&lt;br /&gt;
# Simulate pi&lt;br /&gt;
simulation = 10000&lt;br /&gt;
distances = np.array([pow(uniform(0, 1), 2) + pow(uniform(0, 1), 2) for i in range(simulation)])&lt;br /&gt;
pi = count_nonzero(distances &amp;lt; 1) / simulation * 4&lt;br /&gt;
pi&lt;br /&gt;
&lt;br /&gt;
# Simulate e&lt;br /&gt;
total = 10000&lt;br /&gt;
ranges = [[x/total, (x+1)/total] for x in range(total)]&lt;br /&gt;
sample_space = [uniform(0, 1) for x in range(total)]&lt;br /&gt;
index = []&lt;br /&gt;
for i in range(total):&lt;br /&gt;
    if any(ranges[i][0] &amp;lt;= point &amp;lt;= ranges[i][1] for point in sample_space):&lt;br /&gt;
        index.append(i)&lt;br /&gt;
e = total / (total - len(set(index)))&lt;br /&gt;
e&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Feb 24, 2017 &amp;quot;Idiots&amp;quot; (Due 3/3/2017)==&lt;br /&gt;
[[File:Idiots.png|thumbnail]]&lt;br /&gt;
* Source: Paul Nahin (2000), &amp;quot;Dueling Idiots&amp;quot;. Problem #2: &amp;quot;When Idiots Duel&amp;quot;&lt;br /&gt;
* Game: Idiots A and B decide to duel with a gun in the following way: They will insert a single bullet into the gun&#039;s cylinder. The cylinder has a total of 6 slots. Idiot A will spin the cylinder and shoot at B. If the gun doesn&#039;t fire, then A will give the gun to B, who will spin the cylinder and then shoot at A. This back-and-forth duel will continue until one fool shoots (and kills) the other. &lt;br /&gt;
* Questions: (1) what is the probability that A will win (and B dies); (2) What is the average number of trigger pulls before someone dies?&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
| &lt;br /&gt;
* By Roy&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
gun &amp;lt;-c(1,0,0,0,0,0)&lt;br /&gt;
gun2 &amp;lt;- c(1,0,0,0,0,0)&lt;br /&gt;
deadman &amp;lt;-0 ;idiot1win &amp;lt;-0; idiot2win &amp;lt;- 0; nooned &amp;lt;- 0; total &amp;lt;- 0&lt;br /&gt;
 while(deadman &amp;lt; 1000){&lt;br /&gt;
  idiot1shot &amp;lt;-sample(gun)&lt;br /&gt;
  if (length(which(idiot1shot[1] == 1))){&lt;br /&gt;
    deadman &amp;lt;- deadman + 1&lt;br /&gt;
    idiot1win &amp;lt;- idiot1win + 1&lt;br /&gt;
  }else{&lt;br /&gt;
      idiot2shot &amp;lt;-sample(gun2)&lt;br /&gt;
      if (length(which(idiot2shot[1] == 1))){&lt;br /&gt;
        deadman &amp;lt;- deadman +1&lt;br /&gt;
        idiot2win &amp;lt;- idiot2win +1&lt;br /&gt;
      } else {nooned &amp;lt;- nooned + 1}&lt;br /&gt;
  }&lt;br /&gt;
  total &amp;lt;- total + 1&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
deadman&lt;br /&gt;
idiot1win&lt;br /&gt;
idiot2win&lt;br /&gt;
nooned&lt;br /&gt;
total&lt;br /&gt;
&lt;br /&gt;
p &amp;lt;- idiot1win/1000&lt;br /&gt;
p*100&lt;br /&gt;
takes2kill &amp;lt;- deadman/total&lt;br /&gt;
takes2kill*100&lt;br /&gt;
idiot1shot&lt;br /&gt;
idiot2shot&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* By Weigang&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
rounds &amp;lt;- sapply(1:1000, function(x) {&lt;br /&gt;
  alive &amp;lt;- 1;&lt;br /&gt;
  round &amp;lt;- 0;&lt;br /&gt;
  while(alive == 1){&lt;br /&gt;
    spin &amp;lt;- sample(c(0,0,0,0,0,1));&lt;br /&gt;
    round &amp;lt;- round + 1;&lt;br /&gt;
    if(spin[1] == 1) { alive &amp;lt;- 0 }&lt;br /&gt;
  }&lt;br /&gt;
  return(round)&lt;br /&gt;
})&lt;br /&gt;
prob.a.live &amp;lt;- length(which(rounds %% 2 == 0))&lt;br /&gt;
barplot(table(rounds)/1000, xlab=&amp;quot;num rounds&amp;quot;, ylab=&amp;quot;Prob&amp;quot;, main = &amp;quot;Sudden Death with a 6-slot gun (sim=1000)&amp;quot;, las=1)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Feb 17, 2017 &amp;quot;Birthday&amp;quot; (Due 2/24/2017)==&lt;br /&gt;
[[File:B-day.png|thumbnail]]&lt;br /&gt;
* Problem: What is the probability NONE of the N people in a room sharing a birthday?&lt;br /&gt;
# Randomly select N individuals and record their B-days&lt;br /&gt;
# Count the B-days NOT shared by ANY two individuals&lt;br /&gt;
# Repeat (for each N) 1000 times, and obtain probability by averaging the previous counts (i.e., divided by 1000)&lt;br /&gt;
# Vary N from 10 to 100, increment by 10&lt;br /&gt;
# Plot probability of no-shared B-Day (Y-axis) versus N (x-axis), with either a stripchart or boxplot, or both&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
| &lt;br /&gt;
* By Roy&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
N &amp;lt;- 0&lt;br /&gt;
samples &amp;lt;- 0;&lt;br /&gt;
prob.nodub &amp;lt;-0;&lt;br /&gt;
for(j in 10:100){&lt;br /&gt;
  counting.no.dups &amp;lt;- 0;&lt;br /&gt;
  test &amp;lt;- for(i in 1:1000){&lt;br /&gt;
  bdays &amp;lt;- sample(seq(as.Date(&#039;1990/01/01&#039;), as.Date(&#039;1990/12/31&#039;), by=&amp;quot;day&amp;quot;), N, replace=T)&lt;br /&gt;
  dups &amp;lt;- duplicated(bdays, incomparables = FALSE)&lt;br /&gt;
  ch &amp;lt;-length(which(dups == TRUE))&lt;br /&gt;
    if(ch==0){&lt;br /&gt;
      counting.no.dups &amp;lt;- counting.no.dups +1&lt;br /&gt;
    }&lt;br /&gt;
    fine &amp;lt;- (counting.no.dups/1000)&lt;br /&gt;
  }&lt;br /&gt;
  print(N)&lt;br /&gt;
  print(fine)&lt;br /&gt;
  N &amp;lt;- N + 1&lt;br /&gt;
  samples[[j]] &amp;lt;- N&lt;br /&gt;
  prob.nodub[[j]] &amp;lt;- fine&lt;br /&gt;
}&lt;br /&gt;
plot(samples,prob.nodub, main=&amp;quot;Birthday Simulation&amp;quot;, xlab = &amp;quot;Sample Size&amp;quot;, ylab = &amp;quot;Probability of no Duplicates&amp;quot;, las =1)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* By weigang&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
days &amp;lt;- 1:365;&lt;br /&gt;
&lt;br /&gt;
find.overlap &amp;lt;- function(x) { return(length(which(table(x)&amp;gt;1))) }&lt;br /&gt;
&lt;br /&gt;
output &amp;lt;- sapply(1:100, function(x) { # num of people in the room&lt;br /&gt;
  ct.no.overlap &amp;lt;- 0;&lt;br /&gt;
  for (k in 1:100) {&lt;br /&gt;
    bdays &amp;lt;- sample(days, x, replace = T);&lt;br /&gt;
    ct &amp;lt;- find.overlap(bdays);&lt;br /&gt;
    if (!ct) { ct.no.overlap &amp;lt;- ct.no.overlap + 1}&lt;br /&gt;
  }&lt;br /&gt;
  return(ct.no.overlap);&lt;br /&gt;
})&lt;br /&gt;
 &lt;br /&gt;
plot(1:100, output/100, xlab=&amp;quot;Group size&amp;quot;, ylab = &amp;quot;Prob (no shared b-day)&amp;quot;, las=1, main=&amp;quot;B-day (sim=100 times)&amp;quot;, type=&amp;quot;l&amp;quot;)&lt;br /&gt;
abline(h=0.5, lwd=2, col=2, lty=2)&lt;br /&gt;
abline(v=seq(0, 100, 5), col=&amp;quot;gray&amp;quot;)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Feb 10, 2017 &amp;quot;Dating&amp;quot; (Valentine&#039;s Day Special; Due 2/17/2017)==&lt;br /&gt;
[[File:Dating.png|thumbnail]]&lt;br /&gt;
* Source: Paul Nahin (2008), &amp;quot;Digital Dice&amp;quot;. Problem #20: &amp;quot;An Optimal Stopping Problem&amp;quot;&lt;br /&gt;
* Problem: What is the optimal time point when one should stop dating more people and settle on a mate choice (and live with the decision)&lt;br /&gt;
* Your best strategy is to date an initial sample of N individuals, rejecting all, and marry the next one ranked higher than any of your N individuals. The question is what is the optimal number for N.&lt;br /&gt;
# The problem could be investigated by simulating a pool of 10 individuals, ranked from 1-10 (most desirable being 1) and then take a sample of N&lt;br /&gt;
# You may only date one individual at a time&lt;br /&gt;
# You cannot go back to reach previously rejected candidates&lt;br /&gt;
# Simulate N from 0 to 9 (zero means marrying the first date, a sample size of zero)&lt;br /&gt;
# For each N, obtain the probability of finding the perfect mate (i.e., ranked 1st) by running simulation 1000 times&lt;br /&gt;
# Plot barplot of probability versus sample size N.&lt;br /&gt;
# Expected answer: N=4&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
| &lt;br /&gt;
* By Weigang&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
pick.candidate &amp;lt;- function(min, array) {&lt;br /&gt;
  for (i in 1:length(array)) {&lt;br /&gt;
    if (array[i] &amp;lt; min) {&lt;br /&gt;
      return(array[i])&lt;br /&gt;
    } else {next}&lt;br /&gt;
  }&lt;br /&gt;
  return(0) # No 1. has been sampled and rejected&lt;br /&gt;
}&lt;br /&gt;
candidates &amp;lt;- 1:10;&lt;br /&gt;
output &amp;lt;- sapply(0:9, function(x) {&lt;br /&gt;
  ct &amp;lt;- 0;&lt;br /&gt;
  for(k in 1:1000) {&lt;br /&gt;
    if (x==0) { # no sample, marry the 1st guy &lt;br /&gt;
      sampled &amp;lt;- sample(candidates, 1);&lt;br /&gt;
      if (sampled == 1) {ct &amp;lt;- ct+1}&lt;br /&gt;
    } else {&lt;br /&gt;
      sampled &amp;lt;- sample(candidates, x);&lt;br /&gt;
      not.sampled &amp;lt;- candidates[-sampled];&lt;br /&gt;
      not.sampled &amp;lt;- sample(not.sampled);&lt;br /&gt;
      if (pick.candidate(min =  min(sampled), array = not.sampled) == 1) {ct &amp;lt;- ct+1}&lt;br /&gt;
    }&lt;br /&gt;
  }&lt;br /&gt;
  return(ct);&lt;br /&gt;
})&lt;br /&gt;
barplot(output/1e3, names.arg = 0:9, xlab = &amp;quot;number of sampled dates&amp;quot;, las=1, main = &amp;quot;Optimal stopping for dating (N=10 candidates)&amp;quot;, ylab = &amp;quot;Prob(marrying No.1)&amp;quot;)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Feb 3, 2017 &amp;quot;US Presidents&amp;quot; (Due 2/10/2017)==&lt;br /&gt;
[[File:Sim-presidents.png|thumbnail]]&lt;br /&gt;
* Download [[File:Presidents.txt|thumbnail]]: 1st column is the order, 2nd column is the name, the 3rd column is the year of inauguration; tab-separated&lt;br /&gt;
* Your job is to create an R, Perl, or Python script called “us-presidents”, which will&lt;br /&gt;
# Read the table&lt;br /&gt;
# Store the original/correct order&lt;br /&gt;
# Shuffle/permute the rows and record the new order&lt;br /&gt;
# Count the number of matching orders&lt;br /&gt;
# Repeat Steps 3-4 for a 1000 times&lt;br /&gt;
# Plot histogram or barplot (better) to show distribution of matching counts&lt;br /&gt;
# Hint: For R, use the sample() function. For Perl, use the rand() function.&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
| &lt;br /&gt;
* By Mei&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
pres.list &amp;lt;- lapply(1:1000, function(x) pres[sample(nrow(pres)),])&lt;br /&gt;
cts &amp;lt;- sapply(pres.list, function(x) {&lt;br /&gt;
  ct.match &amp;lt;- 0;&lt;br /&gt;
  for (i in 1:45){&lt;br /&gt;
    if (pres$order[i] == x[i,1]){&lt;br /&gt;
      #cat(as.character(x[i,2]), x[i,1], &amp;quot;\n&amp;quot;)  #as.character to avoid the factor info&lt;br /&gt;
      ct.match &amp;lt;- ct.match + 1;&lt;br /&gt;
      #cat(ct.match)&lt;br /&gt;
    }&lt;br /&gt;
  }&lt;br /&gt;
  ct.match #return&lt;br /&gt;
})&lt;br /&gt;
barplot(table(cts),xlab = &amp;quot;Number of matches per shuffle&amp;quot;, border = &amp;quot;hotpink&amp;quot;, col = &amp;quot;pink&amp;quot;, ylab = paste(&amp;quot;Frequency total of:&amp;quot;,length(pres.list)), main = &amp;quot;US Presidents&amp;quot;)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* By John&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;python&amp;quot;&amp;gt;&lt;br /&gt;
import pandas as pd&lt;br /&gt;
import numpy as np&lt;br /&gt;
from pandas import DataFrame&lt;br /&gt;
import matplotlib.pyplot as plt&lt;br /&gt;
&lt;br /&gt;
df = pd.read_table(&amp;quot;presidents.txt&amp;quot;, names=[&amp;quot;num&amp;quot;, &amp;quot;name&amp;quot;, &amp;quot;presidency&amp;quot;])&lt;br /&gt;
&lt;br /&gt;
name_list = list(df.name)&lt;br /&gt;
&lt;br /&gt;
# create a list to store matches after each shuffle&lt;br /&gt;
shuffle_record = []&lt;br /&gt;
&lt;br /&gt;
# create a function in which the first argument is the original dataset; second argument is number of shuffles&lt;br /&gt;
def shuffler2(original, n):&lt;br /&gt;
    record = []&lt;br /&gt;
    for i in range(n):&lt;br /&gt;
        num = 0&lt;br /&gt;
        each_shuffle = {}&lt;br /&gt;
        temp = original.reindex(np.random.permutation(original.index)) # do shuffling for each&lt;br /&gt;
        compare = original.num == temp.num&lt;br /&gt;
        matched_df = original.ix[compare[compare == True].index]&lt;br /&gt;
        for i in matched_df.index:&lt;br /&gt;
            each_shuffle[i] = matched_df.name[i]&lt;br /&gt;
        shuffle_record.append(each_shuffle)&lt;br /&gt;
        try:&lt;br /&gt;
            num = compare.value_counts()[1]&lt;br /&gt;
        except:&lt;br /&gt;
            pass&lt;br /&gt;
        record.append(num)&lt;br /&gt;
    return(record)&lt;br /&gt;
result2 = shuffler2(df, 1000)&lt;br /&gt;
&lt;br /&gt;
plt.hist(result2, color=&amp;quot;yellow&amp;quot;)&lt;br /&gt;
plt.title(&amp;quot;Histgram for President Data&amp;quot;)&lt;br /&gt;
plt.show()&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt; &lt;br /&gt;
* By Weigang&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
p &amp;lt;- read.table(&amp;quot;Presidents.txt&amp;quot;, sep=&amp;quot;\t&amp;quot;, header=F)&lt;br /&gt;
colnames(p) &amp;lt;- c(&amp;quot;order&amp;quot;, &amp;quot;name&amp;quot;, &amp;quot;inaug.year&amp;quot;)&lt;br /&gt;
# Use &amp;quot;sapply&amp;quot; or &amp;quot;lapply&amp;quot; for loops: no need to pre-define a vector to store results&lt;br /&gt;
p.sim &amp;lt;- sapply(1:10000, function (x) {&lt;br /&gt;
  length(which(sample(p$order) == p$order))&lt;br /&gt;
  }) &lt;br /&gt;
barplot(table(p.sim)/1e4, las=1)&lt;br /&gt;
p.exp &amp;lt;- rpois(1e4, 1) # draw 10000 Poisson random deviates&lt;br /&gt;
mp &amp;lt;- barplot(table(p.exp)/1e4, las=1, xlab = &amp;quot;Num of matching presidents&amp;quot;) # mid-point on x-axis&lt;br /&gt;
lines(mp[1:7], table(p.sim)/1e4, type=&amp;quot;b&amp;quot;, col=2) # add a line (Poisson-expected) to the barplot (simulated)&lt;br /&gt;
legend(&amp;quot;topright&amp;quot;, c(&amp;quot;Simulated&amp;quot;, &amp;quot;Poisson expectation&amp;quot;), col=1:2, lty=1)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Monte_Carlo_Club&amp;diff=6706</id>
		<title>Monte Carlo Club</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Monte_Carlo_Club&amp;diff=6706"/>
		<updated>2026-07-03T19:38:40Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Simulations: Non-adaptive selection */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;__FORCETOC__&lt;br /&gt;
=Mcirobiome &amp;amp; tick pop gen: Ecology &amp;amp; Evolution perspectives (Spring &amp;amp; Summer 2026)=&lt;br /&gt;
==Simulations: Non-adaptive selection==&lt;br /&gt;
* Heterosis: maintenance of genomic diversity by heterozygote advantages&lt;br /&gt;
* Stability of the food web&lt;br /&gt;
* May threshold: Ecological stability: simulations by Brian : https://github.com/aclaAPCSA/ACLAatQiuLab/blob/main/mayFunction.ipynb&lt;br /&gt;
* Prey-predator enrichment&lt;br /&gt;
&lt;br /&gt;
==Simulations: artificial immune systems (AIS)==&lt;br /&gt;
* r-contiguous symbols rule for antigen recognition: https://www.pnas.org/doi/10.1073/pnas.90.5.1691&lt;br /&gt;
* Negative selection: weed out self-antigens, http://dpnm.postech.ac.kr/ANM/resources/papers/09_Nature/hofmeyr_forrest.pdf&lt;br /&gt;
* Clonal selection algorithm for learning: https://ieeexplore.ieee.org/abstract/document/1011539&lt;br /&gt;
&lt;br /&gt;
==Pan-genome graphs:  an efficient computational representation of whole-genomic variations==&lt;br /&gt;
* A summary (including data standards &amp;amp; libraries): https://pangenome.github.io/&lt;br /&gt;
* A review paper: https://www.nature.com/articles/s41588-024-02029-6&lt;br /&gt;
* A paper (Huamn &amp;amp; Drosophila models): https://www.nature.com/articles/s41587-023-01793-w&lt;br /&gt;
* SARS-CoV-2 application: whole-genome alignment with Twilight &amp;amp; compression by [https://github.com/TurakhiaLab/panman panMAN]: https://www.amd.com/en/developer/resources/technical-articles/2026/ucsd-researchers-use-amd-gpus-to-build-the-worlds-large.html&lt;br /&gt;
* An application: PHI, for pangenome-based genotyping: https://github.com/at-cg/PHI&lt;br /&gt;
* Build bacterial pangenome with Panaroo: https://link.springer.com/article/10.1186/s13059-020-02090-4&lt;br /&gt;
* Phylogenetic compression of bacterial genome data: MiniPhy &amp;amp; PanMan algorithms; https://www.nature.com/articles/s41579-026-01303-0&lt;br /&gt;
&lt;br /&gt;
==Tick genomics &amp;amp; native immunity==&lt;br /&gt;
===Papers found by Hui===&lt;br /&gt;
* Abbas, M. N., Jmel, M. A., Mekki, I., Dijkgraaf, I., &amp;amp; Kotsyfakis, M. (2023). Recent Advances in Tick Antigen Discovery and Anti-Tick Vaccine Development. International journal of molecular sciences, 24(5), 4969. https://doi.org/10.3390/ijms24054969. I found this review article to be the most complete and well-rounded one. It summarizes recent progress in anti-tick vaccine development, especially antigen-based strategies that target tick feeding, reproduction, attachment, and pathogen transmission. The authors also explain how modern approaches such as genomics, transcriptomics, proteomics, and reverse vaccinology have helped identify new vaccine candidates from tick tissues including the salivary glands, midgut, eggs, Malpighian tubules, and cement. The review covers several major antigens and vaccine targets, including Bm86/Bm95, ferritin, aquaporins, subolesin, Salp15, metalloproteases, serpins, and the recent 19ISP mRNA vaccine targeting Ixodes scapularis salivary proteins.&lt;br /&gt;
* Allen J.R., Humphreys S.J. Immunisation of guinea pigs and cattle against ticks. Nature. 1979;280:491–493. doi: 10.1038/280491a0. https://www-nature-com.proxy.wexler.hunter.cuny.edu/articles/280491a0. The first review article referred to this early study, so I thought it was important to include. This paper is an early proof-of-concept showing that hosts can be immunized against ticks using antigens extracted from the internal organs of partially fed ticks, especially the midgut and reproductive organs. It helped establish the idea that internal tick antigens, not only salivary proteins, could be useful vaccine targets.&lt;br /&gt;
* Hart, T. M., Cui, Y., Telford, S. R., Marín-López, A., Calloway, K., Dai, Y., Matias, J., DePonte, K., Jaycox, J., DeBlasio, M., Hoornstra, D., Belperron, A. A., Cibichakravarthy, B., Johnson, E. E., Alameh, M. G., Dwivedi, G., Hovius, J. W. R., Bockenstedt, L. K., Weissman, D., Ring, A. M., … Fikrig, E. (2025). Tick feeding or vaccination with tick antigens elicits immunity to the Ixodes scapularis exoproteome in guinea pigs and humans. Science translational medicine, 17(791), eads9207. https://doi.org/10.1126/scitranslmed.ads9207. This paper developed a high-throughput yeast display platform called IscREAM, which can screen antibody responses against more than 3,000 predicted extracellular proteins from Ixodes scapularis. Using vaccinated guinea pigs, Lyme disease patients, and a naturally tick-resistant individual, the authors identified 199 immunogenic tick antigens. They also showed that both salivary and cement proteins may contribute to acquired tick resistance.&lt;br /&gt;
* Narasimhan, Sukanya et al. “Immunity against Ixodes scapularis salivary proteins expressed within 24 hours of attachment thwarts tick feeding and impairs Borrelia transmission.” PloS one vol. 2,5 e451. 16 May. 2007, https://journals.plos.org/plosone/article?id=10.1371%2Fjournal.pone.0000451. This article focuses on early Ixodes scapularis salivary proteins that are important for tick feeding and Borrelia burgdorferi transmission. The authors show that the tick salivary gland transcriptome and proteome change during feeding, especially between the first 24 hours of attachment and later feeding stages. Importantly, they found that guinea pigs repeatedly exposed only to the first 24 hours of tick feeding developed acquired tick immunity, shown by inflammation at the bite site, rapid tick rejection, reduced engorgement, and impaired Borrelia transmission.&lt;br /&gt;
&lt;br /&gt;
===Tick genomics===&lt;br /&gt;
* Mini-review: [https://www.frontiersin.org/journals/cellular-and-infection-microbiology/articles/10.3389/fcimb.2018.00176/full Kitsou &amp;amp; Pal (2018).] Ixodes immune responses against Lyme disease pathogens. Font Cell Infect Microb.&lt;br /&gt;
* The IMD pathway: [https://journals.asm.org/doi/10.1128/mbio.00703-22 Sidak-Loftis et al (2022)]; mBio&lt;br /&gt;
* The PERK pathway: [https://journals.asm.org/doi/10.1128/msphere.00321-23 Rosche et al (2023)]; mSphere&lt;br /&gt;
* Tick microbiome: [https://journals.asm.org/doi/10.1128/spectrum.01404-23 Paulson et al (2023);] Microbio Spectrum&lt;br /&gt;
* The tick genome paper: [https://www.nature.com/articles/s41588-022-01275-w De et al (2023).] Nature Genetics&lt;br /&gt;
* Tcik genome resource:&lt;br /&gt;
&lt;br /&gt;
==Microbiome &amp;amp; Evolution of Complexity: the Holobionts/Gia hypothesis vs Non-adaptive selection==&lt;br /&gt;
* [https://press.uchicago.edu/ucp/books/book/chicago/N/bo244157692.html Non-adaptive Selection], by John Damuth &amp;amp; Lev Ginzburg&lt;br /&gt;
* C.H. Lean, W.F. Doolittle, &amp;amp; J.P. Bielawski, Community-level evolutionary processes: Linking community genetics with &#039;&#039;&#039;replicator-interactor theory&#039;&#039;&#039;, Proc. Natl. Acad. Sci. U.S.A. 119 (46) e2202538119, https://doi.org/10.1073/pnas.2202538119 (2022).&lt;br /&gt;
* Guilhem Doulcier, Amaury Lambert, Silvia De Monte, Paul B Rainey (2020) Eco-evolutionary dynamics of nested Darwinian populations and the emergence of community-level heredity eLife 9:e53433. https://doi.org/10.7554/eLife.53433 &lt;br /&gt;
* Daniel W. McShea, Steve C. Wang, Robert N. Brandon, A quantitative formulation of biology&#039;s first law, Evolution, Volume 73, Issue 6, 1 June 2019, Pages 1101–1115, https://doi.org/10.1111/evo.13735 (the &amp;quot;&#039;&#039;&#039;The zero-force evolutionary law (ZFEL)&#039;&#039;&#039;&amp;quot;)&lt;br /&gt;
* J. Madhusoodanan, Do hosts and their microbes evolve as a unit?, Proc. Natl. Acad. Sci. U.S.A. 116 (29) 14391-14394, https://doi.org/10.1073/pnas.1908139116 (2019).&lt;br /&gt;
* Douglas AE, Werren JH. 2016.Holes in the Hologenome: Why Host-Microbe Symbioses Are Not Holobionts. mBio7:10.1128/mbio.02099-15.https://doi.org/10.1128/mbio.02099-15&lt;br /&gt;
* Ilana Zilber-Rosenberg, Eugene Rosenberg, Role of microorganisms in the evolution of animals and plants: &#039;&#039;&#039;the hologenome theory&#039;&#039;&#039; of evolution, FEMS Microbiology Reviews, Volume 32, Issue 5, August 2008, Pages 723–735, https://doi.org/10.1111/j.1574-6976.2008.00123.x&lt;br /&gt;
* Seth R Bordenstein &amp;amp; Kevin R Theis (2015). Host Biology in Light of the Microbiome: Ten Principles of Holobionts and Hologenomes. PLoS Biology. https://doi.org/10.1371/journal.pbio.1002226&lt;br /&gt;
&lt;br /&gt;
=Coevolution, Misinformation, and Sex (Summer &amp;amp; Fall 2025)=&lt;br /&gt;
==Schedule &amp;amp; Notes==&lt;br /&gt;
* Orientation: 6/19/2025. Brian, Ryan, Emma, Sebastian, Will&lt;br /&gt;
* Week 1. (June 22, 2025, Monday). &lt;br /&gt;
** Presentation: Network evolution and RNA world (by Emma)&lt;br /&gt;
** Presentation: Review last season&#039;s Python work (by Sebastian)&lt;br /&gt;
** Presentation: Python Automata class (by Will)&lt;br /&gt;
** Presentation: Automata &amp;amp; biological applications (by Brandon)&lt;br /&gt;
** Presentation: k-ras data analysis, normalized to DMSO with fold change; linear model; clustering by heatmap (Ahava)&lt;br /&gt;
* Week 2. (June 30, 2025, Monday) &lt;br /&gt;
** Sebastian: Shannon entropy. To do: simulate the small-sample size effect (reduction of uncertainty)&lt;br /&gt;
** Ryan: genetic code &amp;amp; 6-frame translation. To do: 6-frame translation of ospC sequence&lt;br /&gt;
** Brian and Will: Miller Appendix F, Test 1. Adaptation to a fixed string (with genetic algorithm). To do: Test 2. pattern recognition&lt;br /&gt;
** Brandon: k-ras data analysis, gene set enrichment analysis. To do: pick random genes from clusters to validate the results of linear models&lt;br /&gt;
* Week 3. (July 7, 2025)&lt;br /&gt;
** Ryan: done with 6-frame translation &amp;amp; visualization code&lt;br /&gt;
** Sebastian: Simulated entropy given sample size, showing the loss of uncertainty with small sample sizes&lt;br /&gt;
** Brian: done with Miller Appendix F2, evolving automata for pattern recognition&lt;br /&gt;
** Brandon, Ahava, and Lenny: k-ras data heatmaps, preliminary GSEA results, and interaction plots&lt;br /&gt;
* Week 4 (July 14, 2025)&lt;br /&gt;
** Ryan and Sebastian: Mutual info implemented&lt;br /&gt;
** Brian: Implement F3, co-evolving problem set (to avoid over-fitting)&lt;br /&gt;
** Brandon: an automaton model of two-host sensing; test generalist vs specialist; test cost of complexity&lt;br /&gt;
* Week 5 (July 21, 2025)&lt;br /&gt;
** Ryan and Sebastian: Simulate correlated counts; constructed function for mutual information (MI) and D-print&lt;br /&gt;
** Brandon &amp;amp; Brian: evolve two-host sensing; implement tradeoff under the constraint of finite number of states&lt;br /&gt;
* Week 6 (Tuesday, July 29, 2025)&lt;br /&gt;
** Ryan &amp;amp; Sebastian: both finished functions for MIR and bootstrap test&lt;br /&gt;
** Next task: implement relative conditional entropy (RCE) as an asymmetric measure of information content between two loci&lt;br /&gt;
* Week 7 (break)&lt;br /&gt;
* Week 8 (Wed, Aug 13, 2025)&lt;br /&gt;
** Ryan &amp;amp; Sebastian: RCE on Lyme data with bootstrap tests; Quantify genetic differentiation between two populations: Fst (hz, pi) &amp;amp; Jensen-Shannon divergence&lt;br /&gt;
** Brian: automata model for generalists and specialist&lt;br /&gt;
** Brandon: hybrid models with two loci, one under immune selection, another under adaptive selection. Results similar to immune-selection model alone&lt;br /&gt;
* Week 9 (Friday, Aug 21, 2025)&lt;br /&gt;
** End-of-summer party&lt;br /&gt;
** Bootstrap on RCE, by break linkage&lt;br /&gt;
** Agricultural optimization problem&lt;br /&gt;
** Quarto book/website to summarize all work (code and visualization)&lt;br /&gt;
&lt;br /&gt;
== Reading list==&lt;br /&gt;
* &#039;&#039;&#039;Coevolutionary automata&#039;&#039;&#039; John Miller (2023). [https://www.sfipress.org/books/ex-machina Ex Machina]. [https://github.com/SantaFeInstitute/ExMachina?tab=readme-ov-file PDF book on GitHub]&lt;br /&gt;
* &#039;&#039;&#039;Network evolution&#039;&#039;&#039; (automata model): Phylip Ball, Nautilus (2025): [https://nautil.us/the-strange-inevitability-of-evolution-235189/ The Strange Inevitability of Evolution] &lt;br /&gt;
* David P. Barash (2025). [https://nautil.us/evolution-is-really-not-that-into-sex-237397/ Evolution Is Really Not That Into Sex] &lt;br /&gt;
* Hartfield &amp;amp; Keighterley (2012). &amp;quot;Current hypotheses for the evolution of sex and recombination&amp;quot;.https://pubmed.ncbi.nlm.nih.gov/22691203/&lt;br /&gt;
* wikipedia chapter on mutual information: https://en.wikipedia.org/wiki/Mutual_information&lt;br /&gt;
* David MacKay (2004). Chapter 19 &amp;quot;Why have Sex? Information Acquisition and Evolution&amp;quot;&lt;br /&gt;
* Chanda et al (2020). &amp;quot;Information Theory in Computational Biology: Where We Stand Today.&amp;quot; Entropy 2020, 22, 627. https://doi.org/10.3390/e22060627&lt;br /&gt;
&lt;br /&gt;
==Goals and Applications==&lt;br /&gt;
* network evolution and automata model of evolution (e.g., two-host adaptation of Lyme pathogens)&lt;br /&gt;
* genome-wide linkage: multivariate/total correlation&lt;br /&gt;
* genetic costs &amp;amp; benefits of recombination&lt;br /&gt;
&lt;br /&gt;
==Tools &amp;amp; Tasks==&lt;br /&gt;
* Replicate/Implement simulations in MacKay (204) Chapter 19 &lt;br /&gt;
* Python package &amp;lt;code&amp;gt;dit&amp;lt;/code&amp;gt; (discrete info theory): https://dit.readthedocs.io/en/latest/measures/shannon.html#mutual-information&lt;br /&gt;
* R packages for information theory:&lt;br /&gt;
** Distances: philentropy, https://github.com/drostlab/philentropy; https://cran.r-project.org/web/packages/philentropy/vignettes/Information_Theory.html&lt;br /&gt;
** Package: entropy&lt;br /&gt;
** Package: infotheo&lt;br /&gt;
&lt;br /&gt;
==Applications==&lt;br /&gt;
* genome-wide linkage: multivariate/total correlation&lt;br /&gt;
&lt;br /&gt;
=Bb pangenomics project (Fall 2024, with Brian &amp;amp; students)=&lt;br /&gt;
==Reading list==&lt;br /&gt;
* Pop-gen Theory: Baumdicker &amp;amp; Kupczok (2023), &amp;quot;Tackling the Pangenome Dilemma Requires the Concerted Analysis of Multiple Population Genetic Processes&amp;quot;, GBE. https://academic.oup.com/gbe/article/15/5/evad067/7137407?login=false &lt;br /&gt;
* Akther et al (2024). &amp;quot;Natural selection and recombination at host-interacting lipoprotein loci drive genome diversification of Lyme disease and related bacteria&amp;quot; [https://journals.asm.org/doi/10.1128/mbio.01749-24 mBio 0:e01749-24.]&lt;br /&gt;
* Tonkin-Hill  et al (2023). &amp;quot;Robust analysis of prokaryotic pangenome gene gain and loss rates with Panstripe&amp;quot;. [https://genome.cshlp.org/content/33/1/129.full?sid=07484132-d2b7-4fa7-bde0-8d182c64929b Genome Research]&lt;br /&gt;
* Gautreau et al (2020). &amp;quot;PPanGGOLiN: Depicting microbial diversity via a partitioned pangenome graph&amp;quot;. [https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1007732 PLoS Comp Bio.]&lt;br /&gt;
&lt;br /&gt;
==Project Description==&lt;br /&gt;
* Project Goals&lt;br /&gt;
# Species-specific gene gains &amp;amp; losses&lt;br /&gt;
# Correlation with phenotypes (e.g., geography, tick &amp;amp; host association, virulence &amp;amp; pathogenicity)&lt;br /&gt;
&lt;br /&gt;
* Project outline&lt;br /&gt;
# Phase I. Run ppangglin pipeline for 78 Bb main chromosomes: obtain gene presence/absence&lt;br /&gt;
# Phase II. Run phylogenetic reconstruction: obtain branch-specific gene gains and losses&lt;br /&gt;
# Phase III. Run linear regression: obtain gene accumulation curves&lt;br /&gt;
# Phase IV. Plasmid-borne genes&lt;br /&gt;
&lt;br /&gt;
==Preliminary results==&lt;br /&gt;
* PPanGGoLiN output 1: [https://borreliabase.org/~wgqiu/Ushaped_plot.html U-shaped plot for 47 Bb genomes]&lt;br /&gt;
* PPanGGoLiN output 2: [https://borreliabase.org/~wgqiu/tile_plot.html Heatmap of the presence &amp;amp; absence of genes]&lt;br /&gt;
&lt;br /&gt;
=Season VII. Spring &amp;amp; Summer 2024=&lt;br /&gt;
==Monte Carlo Summer Coding Camp==&lt;br /&gt;
* Github repository: https://github.com/weigangq/MonteCarlo&lt;br /&gt;
* For each problem,&lt;br /&gt;
# Step 1. Students will first learn the use of simulator functions (in R and Python)&lt;br /&gt;
# Step 2. Students will write out pseudo code &lt;br /&gt;
# Step 3. Students will implement in R and Python&lt;br /&gt;
# Step 4. Students will make a summary plot&lt;br /&gt;
# Code will be posted on a github as a permanent citable record&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Caption text&lt;br /&gt;
|-&lt;br /&gt;
! Problem !! Simulator !! Pseudo-code !! Data Structure || Expected figure/output !! Learning goal&lt;br /&gt;
|-&lt;br /&gt;
| Birthday: what is the probability of having at least one shared birthday in a class of N students? || uniform samplers. R: &amp;lt;code&amp;gt;sample(); Python: &amp;lt;code&amp;gt;numpy.random.uniform()  || ?? || list/vector || XY plot: X-axis: class size; Y-axis: prob of at least one shared birthday || simulation as experiments; sampling w/ and w/o replacements &lt;br /&gt;
|-&lt;br /&gt;
| Gene frequency: Frequentist approach: maximum likelihood (MLE), Prob(k given p &amp;amp; N): What is the sampling error of a gene frequency given that k alleles are found in a sample of N haploid individuals? || binomial sampler || ?? || ? || mean &amp;amp; variance/standard error || sampling error; maximum likelihood&lt;br /&gt;
|-&lt;br /&gt;
| Gene frequency: Bayesian approach, Prob(p given k &amp;amp; N): what is the gene frequency given k alleles are found in a sample of N haploid individuals || beta sampler || ?? || ? || distribution of posterior probs given success=1,10,100 and failure=9,90,900  || prior &amp;amp; posterior probabilities; Bayes rule; updating rules for beta function  &lt;br /&gt;
|-&lt;br /&gt;
| Monte Hall problem: Should you switch your pick after the host reveals one of the three doors, behind only one of which is a car? || Uniform sampler || ?? || ?? || ?? || conditional/Bayesian prob by simulation&lt;br /&gt;
|-&lt;br /&gt;
| Blending inheritance: why would evolution grind to a halt if genes blend and children are &amp;quot;averages&amp;quot; of two parents? || Uniform sampler || input: one parental populations of N=1000 individuals, each with equal probability of 0&#039;s and 1&#039;s;  a child is produced by randomly selecting a mom and dad and averaging their values; repeat for G generations. Output: genetic diversity in each generation (sd) || ?? || ?? || genetic diversity&lt;br /&gt;
|-&lt;br /&gt;
| Mendenlian inheritance: How would genetic diversity change with random mating?  || Uniform sampler || input: one parental (diploid) populations of N=1000, with two homozygotes &amp;quot;AA&amp;quot; (e.g., 900) and &amp;quot;GG&amp;quot; (e.g., 100);  a child is produced by randomly selecting a mom and dad, followed by another random selection of gametes (A or G from each parent).  Repeat for G generations. Output: genetic diversity in each generation || ?? || ?? || heterozygosity = 1 - sum(p^2)&lt;br /&gt;
|-&lt;br /&gt;
| Genetic drift  || Binomial sampler || Start with a haploid population of 100 individuals with &amp;quot;A&amp;quot; and another 100 individual with &amp;quot;G&amp;quot;. Reproduce for 500 generations while maintaining the total population size of N=200. Track the frequency of the &amp;quot;A&amp;quot; allele, Repeat by using 10X number of individuals || ?? || ?? || &lt;br /&gt;
|-&lt;br /&gt;
| Simulate a genome  || Uniform sampler || The SARS-CoV-2 genome has ~29K bases that is 70% AT. Simulate the genome with the same AT% || ?? || ?? || &lt;br /&gt;
|-&lt;br /&gt;
| Central Dogma  || Uniform sampler || Simulate DNA duplication, Transcription, and 6-frame translation || ?? || ?? || &lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Summer Projects &amp;amp; Readings==&lt;br /&gt;
* Borrelia evolution (Brandon, Saymon, Brian, Che)&lt;br /&gt;
** Nanopore sequencing, evolutionary analysis&lt;br /&gt;
** Co-evolution simulations &amp;amp; modeling (Brandon)&lt;br /&gt;
** ospC multi-epitope model &lt;br /&gt;
* BorreliaBase (Lia &amp;amp; Yozen): &lt;br /&gt;
** Shiny app deployment (done)&lt;br /&gt;
** code refactoring; &lt;br /&gt;
** variable-plasmid new genome data&lt;br /&gt;
** new visualizations (treemap; collapsable trees); &lt;br /&gt;
** gtf files for downloading&lt;br /&gt;
* Miscellaneous collaborative questions&lt;br /&gt;
** HIV intra-host evolution (Dr Evering)&lt;br /&gt;
** scRNA (Dr Melendaz)&lt;br /&gt;
* Readings: Estimating gene frequencies &amp;amp; testing natural selection in bacterial and viral samples&lt;br /&gt;
** [https://journals.asm.org/doi/10.1128/iai.00244-23?url_ver=Z39.88-2003#supplementary-materials Schifflett et al (2023). OpsC alleles in mammals]. To do: Download data set and replicate the figures (including diversity indices) (Laziz)&lt;br /&gt;
** [https://academic.oup.com/pnasnexus/article/2/8/pgad234/7236418#419492093 Combs et al (2023). OspC frequencies in birds and other hosts.] To do: Supplemental Table S1 &amp;amp; S3&lt;br /&gt;
** [https://elifesciences.org/articles/35962 McCrone et al. Influenza within-host variants]. To do: Download &amp;amp; run model code from github. The presence/absence model, compare with the previous tick model (see below)&lt;br /&gt;
** A method for estimating allele frequencies in bacterial samples: [https://academic.oup.com/genetics/article/158/3/1383/6049706 Anderson &amp;amp; Scheet (2001).]&lt;br /&gt;
** [https://www.mdpi.com/1999-4915/13/7/1216 Zhu et al (2021). A model to estimate viral fitness.] To do: implement the github code&lt;br /&gt;
&lt;br /&gt;
==Spring schedule==&lt;br /&gt;
* Week 2 (Feb 1)&lt;br /&gt;
** Presentation: Wayne (HIV tropism). Data: gp120 seqs. Code: from Arezoo&lt;br /&gt;
** Presentation: Aura &amp;amp; Michelle (Single-cell sequencing on endometrial carcinoma&lt;br /&gt;
** Presentation: Susmita &amp;amp; Lily (Borrelia antigens). Data: variability analysis of chosen antigens&lt;br /&gt;
&lt;br /&gt;
* Week 1 (Jan 25)&lt;br /&gt;
** Presentation: Esther (GBS by GEMMA)&lt;br /&gt;
** Project update: Borrelia transcriptomics visualization &amp;amp; Shiny App&lt;br /&gt;
** Presentation: Liann (Borrelia replication &amp;amp; polyploidy)&lt;br /&gt;
** Project update: Arezoo (dummy data set II, phylogenetic random)&lt;br /&gt;
** An R package to read Prism data: [https://cran.r-project.org/web/packages/pzfx/vignettes/pzfx.html The pzfx package]&lt;br /&gt;
&lt;br /&gt;
==Projects &amp;amp; Readings==&lt;br /&gt;
* Borreliella genomics: &lt;br /&gt;
** Genome &amp;amp; gene variability: [http://test.borreliabase.org/ BorreliaBase]&lt;br /&gt;
** Transcriptomes by [https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0164165 Arnold et al (2016). PLoS One]&lt;br /&gt;
** Diagnostic antigens by [https://jcm.asm.org/content/57/12/e01142-19 Arumugam et al (2019). J. Clin. Micro.]&lt;br /&gt;
** Plasmid partitioning by [https://www.nature.com/articles/s41467-022-34876-4 Takacs et al (2022). Nat Communications]&lt;br /&gt;
** DbpA variability in European isolates, by [https://pubmed.ncbi.nlm.nih.gov/16530482/ Schulte-Spechtel et al (2006). Int J. Med Microbiol]&lt;br /&gt;
** Non-abundant antigens by [https://journals.asm.org/doi/10.1128/msystems.01087-23?utm_source=literatum&amp;amp;utm_medium=email&amp;amp;utm_campaign=alert-emails Yaş et al (2023). mSystems]&lt;br /&gt;
* Single-cell transcriptomes of cancer cells&lt;br /&gt;
** [https://www.nature.com/articles/s41467-022-33982-7 Ren et al (2022). Nat Comm.] &lt;br /&gt;
** [https://pubmed.ncbi.nlm.nih.gov/35560676/ Yu et al (2022). Cell Prolif]&lt;br /&gt;
** [https://satijalab.org/seurat/articles/pbmc3k_tutorial Tutorial: basic pipeline with Suerat]&lt;br /&gt;
** [https://satijalab.org/seurat/articles/pbmc3k_tutorial Tutorial: annotation with singleR]&lt;br /&gt;
* Protein classifiers&lt;br /&gt;
** Simulated data&lt;br /&gt;
** HIV tropism&lt;br /&gt;
*** Salmi et al (2013). J. Leukocyte Biology. [https://academic.oup.com/jleukbio/article-abstract/93/1/113/6936163?redirectedFrom=fulltext&amp;amp;login=false Macrophage-tropic HIV-1 variants from brain demonstrate alterations in the way gp120 engages both CD4 and CCR5]&lt;br /&gt;
*** Sequences from Gray et al (2011). Retrovirology. [https://retrovirology.biomedcentral.com/articles/10.1186/1742-4690-8-42 Conformational alterations in the CD4 binding cavity of HIV-1 gp120 influencing gp120-CD4 interactions and fusogenicity of HIV-1 envelopes derived from brain and other tissues]&lt;br /&gt;
** Toxin specifications&lt;br /&gt;
* Co-evolution: simulations &amp;amp; applications&lt;br /&gt;
** Co-evolutionary antomata: by John Miller (2022). [https://github.com/SantaFeInstitute/ExMachina Github]&lt;br /&gt;
** Evolving synthetic proteins by [https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10403280/ Yang et al (2023) Science]&lt;br /&gt;
&lt;br /&gt;
=Season VI. Language Models, Automata, and Evolutionary Games (Spring, Summer &amp;amp; Fall 2023)=&lt;br /&gt;
==Week 2. Lab meeting (June 8, 2023)==&lt;br /&gt;
* GBS microbial genomics. Participants: Eamen, Esther&lt;br /&gt;
* Conotoxin project: unsupervised clustering by Snekmer (amino acid recoding + kmer). Participants: Arezoo, Jessy&lt;br /&gt;
* Red Queen evolution: Reinforcement learning implementation of tic-tac-toe: [https://github.com/ShangtongZhang/reinforcement-learning-an-introduction RL codes]. Participants: Brandon&lt;br /&gt;
* Card game: [https://evolution-outreach.biomedcentral.com/articles/10.1186/s12052-015-0039-2 Gibson et al (2015). &amp;quot;The Red Queen’s Race: An Experimental Card Game to Teach Coevolution&amp;quot;]. Participants: all&lt;br /&gt;
==Week 1. Summer kickoff at Rockefeller University (June 1, 2023)==&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;1&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File:Summer-kickoff-RU-2023.jpg&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
==Projects &amp;amp; Papers==&lt;br /&gt;
* Red Queen (RH) evolution through co-evolutionary automata models (CAM model)&lt;br /&gt;
** Github &amp;amp; Free PDF book: https://github.com/SantaFeInstitute/ExMachina&lt;br /&gt;
** With Reinforcement learning (RL) &amp;amp; finite Markov decision processes (MDPs): [http://incompleteideas.net/book/the-book-2nd.html Book in PDF]; Python implementations: [https://github.com/ShangtongZhang/reinforcement-learning-an-introduction Python codes] &lt;br /&gt;
* Identification of toxin specificity through machine learning and protein language models&lt;br /&gt;
** Readings &amp;amp; Github repository: [https://github.com/weigangq/cono-models/tree/main https://]https://github.com/weigangq/cono-models&lt;br /&gt;
* Design of vaccinogen through polyclonal data analysis &amp;amp; antigenic landscape&lt;br /&gt;
** Lyme antigen design with genetic algorithms (GA): https://www.nature.com/articles/s41396-021-01089-4&lt;br /&gt;
** Mutational landscape using MSA protein language models: https://www.sciencedirect.com/science/article/pii/S2211124722018526&lt;br /&gt;
* Generative art contest, hosted by the &#039;&#039;Chance&#039;&#039; magazine of ASA (Am Stat Association)&lt;br /&gt;
** Contest website (requirements, categories, and deadline: Aug 31, 2023): https://community.amstat.org/chancemagazinecontest/home&lt;br /&gt;
** Software tool 1. R package &amp;quot;generativeart&amp;quot;: https://github.com/cutterkom/generativeart&lt;br /&gt;
** Software tool 2. R package &amp;quot;GenerativeR&amp;quot;, [https://github.com/MartinMSPedersen/GenerativeR generative art with R and Shiny] &lt;br /&gt;
** Software tool 3. R package &amp;quot;aRtsy&amp;quot;, [https://github.com/koenderks/aRtsy Generative Art with ggplot2]&lt;br /&gt;
* Reinforcement learning (RL) &amp;amp; finite Markov Decision Processes (MDPs) as a adaptive learning model&lt;br /&gt;
** Code re-implemented in Python: https://github.com/ShangtongZhang/reinforcement-learning-an-introduction&lt;br /&gt;
** Book PDF: http://incompleteideas.net/book/the-book-2nd.html&lt;br /&gt;
&lt;br /&gt;
=Season V. Genes, Memes, and Machines (Spring, Summer &amp;amp; fall 2022)=&lt;br /&gt;
* A journal club to continue the exploration of the link between evolution &amp;amp; learning. &lt;br /&gt;
==Week 1==&lt;br /&gt;
* A primer of information, presented by Winston [[File:Information Theory Primer-Winston.pdf|thumbnail]]&lt;br /&gt;
* ITSNTS theory of selection unit, presented by Anh. [[File:ITSNTS - Anh Pham.pdf|thumbnail]]&lt;br /&gt;
==Week 2==&lt;br /&gt;
* Information theory of Individuality, presented by Anh [[File:Information_theory_of_individual_Anh_Pham.pdf]]&lt;br /&gt;
* Information definition of liveliness, presented by Winston [[File:Aliveness.pdf]]&lt;br /&gt;
==Week 3==&lt;br /&gt;
* Code submissions by Niemah&lt;br /&gt;
[[File:Presidents Shuffle bar plot.png|thumbnail]]&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&#039;bash&#039;&amp;gt;&lt;br /&gt;
library(tidyverse)&lt;br /&gt;
p &amp;lt;- read_tsv(&amp;quot;Presidents.txt&amp;quot;, col_names = FALSE )&lt;br /&gt;
colnames(p) &amp;lt;- c(&amp;quot;order&amp;quot;, &amp;quot;name&amp;quot;, &amp;quot;year.elected&amp;quot;)&lt;br /&gt;
true.order &amp;lt;- p$name&lt;br /&gt;
output &amp;lt;- vector(&amp;quot;double&amp;quot;, length = 10000)&lt;br /&gt;
for(i in 1:10000) {&lt;br /&gt;
  random.order &amp;lt;- sample(p$name)&lt;br /&gt;
  output[[i]] &amp;lt;- sum(true.order == random.order)&lt;br /&gt;
}&lt;br /&gt;
x &amp;lt;- table(output)&lt;br /&gt;
df &amp;lt;- tibble(run = 1:1e4, matches = output)&lt;br /&gt;
df %&amp;gt;% ggplot(aes(x=matches)) + geom_bar(fill = &amp;quot;lightgreen&amp;quot;) + theme_minimal()&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&#039;python&#039;&amp;gt;&lt;br /&gt;
import pandas as pd&lt;br /&gt;
import numpy as np&lt;br /&gt;
import matplotlib.pyplot as plt&lt;br /&gt;
import random&lt;br /&gt;
&lt;br /&gt;
# read file&lt;br /&gt;
file = r&#039;/content/Presidents.txt&#039;&lt;br /&gt;
dp = pd.read_table(&#039;Presidents.txt&#039;, header=None, names=[&#039;order&#039;, &#039;name&#039;, &#039;year.elected&#039;])&lt;br /&gt;
print(dp)&lt;br /&gt;
names_list = list(dp.name)&lt;br /&gt;
print(names_list)&lt;br /&gt;
&lt;br /&gt;
# simulate by 10000 permutations&lt;br /&gt;
output = [] # initialize an empty vector to store the match numbers&lt;br /&gt;
for x in range(10000):&lt;br /&gt;
  names_permuted = np.random.permutation(dp.name)&lt;br /&gt;
  num_match = sum(names_list == names_permuted)&lt;br /&gt;
  output.append(num_match)&lt;br /&gt;
print(output)&lt;br /&gt;
&lt;br /&gt;
# tabulate counts&lt;br /&gt;
table_match = {} # initialize a dict to store counts&lt;br /&gt;
for num in output:&lt;br /&gt;
  if num in table_match:&lt;br /&gt;
    table_match[num] += 1&lt;br /&gt;
  else: # initialize if the number is first seen&lt;br /&gt;
    table_match[num] = 0&lt;br /&gt;
print(table_match)&lt;br /&gt;
&lt;br /&gt;
# plot&lt;br /&gt;
fig = plt.figure()&lt;br /&gt;
ax = fig.add_axes([0,0,1,1])&lt;br /&gt;
x = []&lt;br /&gt;
y = []&lt;br /&gt;
for match, cts in table_match.items():&lt;br /&gt;
    x.append(match)&lt;br /&gt;
    y.append(cts)&lt;br /&gt;
ax.bar(x,y)&lt;br /&gt;
ax.set_title(&amp;quot;Number of President Matches&amp;quot;)&lt;br /&gt;
ax.set_xlabel(&#039;Matches&#039;)&lt;br /&gt;
ax.set_ylabel(&#039;Count&#039;)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Week 4==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
!  Entropy plots (Winston &amp;amp; Anh) !! Fitness plots &lt;br /&gt;
|-&lt;br /&gt;
| [[File:Entropy.png|thumbnail]] || [[File:fit.png|thumbnail]] &lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Week 6==&lt;br /&gt;
* Slides by Anh: [[File:Structure and Mutations of SARS-CoV-2 Spike Protein Anh Pham.pdf|thumbnail]]&lt;br /&gt;
* Slides by Winston: [[File:Binding of Human ACE2 and RBD of Omicron.pdf| thumbnail | Presentation by Winston]]&lt;br /&gt;
==Week 7==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Fitness landscapes of binary strings&lt;br /&gt;
|-&lt;br /&gt;
| &lt;br /&gt;
[[File:Landscape.png|thumbnail| three distributions of fitness of mutations (by Anh &amp;amp; Winston)]]&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Reading list==&lt;br /&gt;
* [https://pages.cs.wisc.edu/~dyer/cs540/handouts/info-theory-primer.pdf A primer for information theory]&lt;br /&gt;
* A somewhat advanced introduction of information theory by [https://sites.santafe.edu/~simon/it.pdf Simon DeDeo (2018)]&lt;br /&gt;
* Data ≠ Information: Scharf (2021). The Ascent of Information: Books, Bits, Genes, Machines, and Life&#039;s Unending Algorithms. [https://www.amazon.com/Ascent-Information-Machines-Unending-Algorithm/dp/0593087240 Amazon link]&lt;br /&gt;
* The &amp;quot;It&#039;s the song not the singer&amp;quot; (ITSNTS) theory of selection unit: [https://pubmed.ncbi.nlm.nih.gov/29581311/ Doolittle &amp;amp; Inkpen (2018).] &amp;quot;Processes and patterns of interaction as units of selection: An introduction to ITSNTS thinking&amp;quot;, PNAS.&lt;br /&gt;
* Evolution is an AI machine: &lt;br /&gt;
** [https://www.oreilly.com/radar/open-endedness-the-last-grand-challenge-youve-never-heard-of/ Stanley, Lehman, and Soros (2017).] &amp;quot;Open-endedness: The last grand challenge you’ve never heard of - While open-endedness could be a force for discovering intelligence, it could also be a component of AI itself.&amp;quot; O&#039;Reily&lt;br /&gt;
** Novelty search algorithm: towards evolutionary complexity. a complexification algorithm that can be integrated into any evolutionary computation (EC)&lt;br /&gt;
** [https://www.oreilly.com/radar/open-endedness-a-new-grand-challenge-for-ai/ O&#039;Reily talk video clip by Ken Stanley (2019)]&lt;br /&gt;
** [http://eplex.cs.ucf.edu/ EPLEX group]&lt;br /&gt;
** [https://nbenko1.github.io/#/ Evolutionary art: picbreeder]&lt;br /&gt;
** [http://eplex.cs.ucf.edu/noveltysearch/userspage/#howtoimplement Implementation: add a vector to characterize behavior (not fitness)]&lt;br /&gt;
** NEAT paper: [https://www.cs.swarthmore.edu/~meeden/cs81/f17/projects/JeffKyle.pdf Novak and Richmond-Crosset (2017)]&lt;br /&gt;
** Evolutionary creativity: [[https://direct.mit.edu/artl/article/26/2/274/93255/The-Surprising-Creativity-of-Digital-Evolution-A|Lehman, Clune, and Misevic (2020). Surprising creativity of digital evolution. Alife]]&lt;br /&gt;
* An algorithmic definition of individuality: [https://pubmed.ncbi.nlm.nih.gov/32212028/ Krakauer et al (2020)]. &amp;quot;The information theory of individuality&amp;quot;. Theory in BioSci&lt;br /&gt;
* Evolution towards complexification: [https://aip.scitation.org/doi/abs/10.1063/1.3643064 Krakauer (2011)]. &amp;quot;Darwinian demons, evolutionary complexity, and information maximization&amp;quot;. Chaos.&lt;br /&gt;
* &amp;quot;Aliveness&amp;quot;: [https://link.springer.com/article/10.1007%2Fs12064-020-00331-5 Kim  et al (2021)]. &amp;quot;Informational architecture across non-living and living collectives&amp;quot;. Theory in BioSci&lt;br /&gt;
* [https://www.pnas.org/content/97/9/4463.long Adami et al (2000). Evolution of biological complexity. PNAS]&lt;br /&gt;
* Evolution-based protein designs &amp;amp; structure-function analysis&lt;br /&gt;
** To increase stability: [https://doi.org/10.1016/j.molcel.2016.06.012 Goldenzweig et al (2021).]&lt;br /&gt;
** Full-spectrum mutation scan: [https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3638900/ Whitehead et al (2012).]&lt;br /&gt;
** Full-spectrum mutation scan of 4 sites GB1 protein, adaptive fitness landscape with higher-order epistasis: [https://elifesciences.org/articles/16965 Wu et al (2016) eLife]&lt;br /&gt;
** Full mutation scan of influenza antigenic drift sites: [https://www.nature.com/articles/s41467-020-15102-5 Wu et al (2020)]; [https://pubmed.ncbi.nlm.nih.gov/34878407/ Wang et al (2021)]&lt;br /&gt;
** A review of SARS-CoV-2 S protein structure: effects of common mutations. [https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8673470/ Mehara &amp;amp; Kepp (2021)] [[File:Structure and Mutations of SARS-CoV-2 Spike Protein Anh Pham.pdf|thumbnail]]&lt;br /&gt;
** Computational modeling of Omicron mutations: [https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8811907/ Kim et al (2022)] [[File:Binding of Human ACE2 and RBD of Omicron.pdf| thumbnail | Presentation by Winston]]&lt;br /&gt;
* Predicting fitness of haplotypes from individual mutations; epistasis&lt;br /&gt;
** [https://journals.plos.org/plospathogens/article?id=10.1371/journal.ppat.1009669 Pedruzzi &amp;amp; Rouzine (2021). PLoS Pathogen.] &amp;quot;An evolution-based high-fidelity method of epistasis measurement: theory and application to influenza&amp;quot;: a method to distinguish covariance by linkage and covariance by epistasis (selection), with &amp;quot;tri-way haplotype frequency&amp;quot;&lt;br /&gt;
** Universal exponential Distribution of Fitness Effect (DFE) of mutations, resistant to linkage and epistasis: [https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1008822 Barlukova &amp;amp; Rouzine (2021). PLoS Comp Biol.]&lt;br /&gt;
** Estimate fitness landscape of gp160, a large HIV variable surface protein: [https://www.pnas.org/doi/full/10.1073/pnas.1717765115 Louie et al (2018).] PNAS. &amp;quot;Fitness landscape of the human immunodeficiency virus envelope protein that is targeted by antibodies&amp;quot;&lt;br /&gt;
** Fitness landscape of Ribozyme: [https://journals.plos.org/plosbiology/article?id=10.1371/journal.pbio.3000300 Bendixsen et al (2019).] PLoS Biology. &amp;quot;Genotype network intersections promote evolutionary innovation&amp;quot;&lt;br /&gt;
** Fitness landscape in Dengue virus, based on ultra-RNA sequencing &amp;amp; growth measurements, indicating numerous deleterious mutations. [https://elifesciences.org/articles/61921 Dolan et al.(2021) eLife]&lt;br /&gt;
** Fitness landscape of SARS-CoV-2: genetic and non-genetic components. [https://academic.oup.com/ve/article/7/2/veab073/6363035 Kepler et al (2021). Virus Evolution.]&lt;br /&gt;
** Predicting mutation fitness of SARS-Cov-2: Obermeyer et al (2022). Analysis of 6.4 million SARS-CoV-2 genomes identifies mutations associated with fitness. [https://www.science.org/doi/10.1126/science.abm1208 Science] &lt;br /&gt;
** Deep mutation scan of SARS-CoV-2 RBD-binding domain: [https://www.science.org/doi/epdf/10.1126/science.abo7896 Starr et al (2022). Science]&lt;br /&gt;
** Song &amp;amp; Zhang (2022). &amp;quot;Unbiased inference of the fitness landscape ruggedness from imprecise fitness estimates&amp;quot;. 3 fitness landscapes: NK, RMF, and polynomial fitness models. [https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9018209/ PubMed link]. [https://github.com/song88180/fitness-landscape-error Github repository]&lt;br /&gt;
&lt;br /&gt;
==Tookits==&lt;br /&gt;
* [https://sites.santafe.edu/~simon/page7/page7.html THOTH: a Python package for Shanon information]&lt;br /&gt;
* [https://github.com/weigangq/cov-db CoV simulator on Github]&lt;br /&gt;
* [https://www.laurentluce.com/posts/solving-mazes-using-python-simple-recursivity-and-a-search/ A maze solver using novelty search (in Python)]&lt;br /&gt;
* [https://github.com/yaricom/goNEAT_NS Github for NEAT-based maze solver]: behavioral vector: ending position:&lt;br /&gt;
** &amp;quot;Therefore, because what is important in a maze is where the solving agent ends, for the maze domain, the behavior of a navigator is defined as its ending position. The novelty metric is then the N-nearest neighbor distance novelty between the ending positions of all known solving agents.&amp;quot;&lt;br /&gt;
* A collection of divergence and diversity search: [https://github.com/DanieleGravina/divergence-and-quality-diversity github]: novelty search, surprise search, curiosity search, quality diversity, etc&lt;br /&gt;
* [https://github.com/CodeReclaimers/neat-python NEAT-Python] &amp;amp; various forks on Github&lt;br /&gt;
* Novelty search by [https://github.com/peter-ch/MultiNEAT multiNEAT], python implementation&lt;br /&gt;
* Visualization of fitness landscape &amp;amp; path with R&lt;br /&gt;
** [https://r-charts.com/correlation/contour-plot/ Basic contour function] &lt;br /&gt;
** [https://plotly.com/r/contour-plots/ Contour map by R plotly]&lt;br /&gt;
* Links for Bayesian nested models&lt;br /&gt;
** with rstan: [https://biologyforfun.wordpress.com/2016/12/08/crossed-and-nested-hierarchical-models-with-stan-and-r/ on WorldPress]&lt;br /&gt;
** [https://www.r-bloggers.com/2016/12/crossed-and-nested-hierarchical-models-with-stan-and-r/ | another example on r-bloggers]&lt;br /&gt;
** [https://www.r-bloggers.com/2016/12/crossed-and-nested-hierarchical-models-with-stan-and-r/ Nested model with categorical predictors]&lt;br /&gt;
&lt;br /&gt;
=Season IV. Classification using Machine Learning (Summer 2021)=&lt;br /&gt;
* Textbook: Aurélien Géron (2019). Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems, 2nd Edition. [https://www.amazon.com/gp/product/1492032646/ref=ppx_yo_dt_b_asin_title_o00_s00?ie=UTF8&amp;amp;psc=1 Amazon link]&lt;br /&gt;
* Set up Python work environment with [https://docs.conda.io/en/latest/miniconda.html mini-conda]&lt;br /&gt;
* We will be using Jupyter-notebook (part of mini-codon installation) to share codes&lt;br /&gt;
==Week 1. MNIST dataset (Chapter 3)==&lt;br /&gt;
# Dataset loading and display (pg 85-87): Jackie, Niemah, and Hannah&lt;br /&gt;
# Binary classifier&lt;br /&gt;
## Cross-validation (pg 89-90): Roman, etc&lt;br /&gt;
## Confusion matrix (pg 90-92): &lt;br /&gt;
## Precision and recall (pg 92-97)&lt;br /&gt;
## ROC curve (pg. 97-100)&lt;br /&gt;
# Multiclass classification (pg 100-108): Brian, etc&lt;br /&gt;
&lt;br /&gt;
==Week 2. K-means clustering (Chapter 9)==&lt;br /&gt;
# 2D simulated dataset&lt;br /&gt;
# Image recognition&lt;br /&gt;
# MNIST dataset&lt;br /&gt;
==Week 3. Exercises (pg.275-276, Chapter 9)==&lt;br /&gt;
# Exercise 10: Facial recognition (Olivetti faces dataset, with k-means)&lt;br /&gt;
# Exercise 11. Facial recognition (semi-supervised learning&lt;br /&gt;
&lt;br /&gt;
=Notes on Origin of Life (Spring 2020)=&lt;br /&gt;
# [https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0224552 Attie et al (2019). Genetic Code optimized as a traveling salesman problem]&lt;br /&gt;
# [https://itsatcuny.org/calendar/self-organizing-systems-and-the-origin-of-life GC Origin of Life Seminar (2/19/2021)]&lt;br /&gt;
# [https://www.cell.com/current-biology/fulltext/S0960-9822(15)00681-8?_returnURL=https%3A%2F%2Flinkinghub.elsevier.com%2Fretrieve%2Fpii%2FS0960982215006818%3Fshowall%3Dtrue Pressman et al (2015). Review: RNA World (ribozyme as origin of genetic code)]&lt;br /&gt;
# A perspective: [https://www.nature.com/articles/s42256-020-00278-8 Miikkulainen, R., Forrest, S. A biological perspective on evolutionary computation. Nat Mach Intell 3, 9–15 (2021)]&lt;br /&gt;
&lt;br /&gt;
=Season III. Summer 2018 (Theme: Evolutionary Computing)=&lt;br /&gt;
==Week 1. Introduction &amp;amp; Motivating Examples==&lt;br /&gt;
# Genetic Arts&lt;br /&gt;
## [http://picbreeder.org Go to picbreeker website] &amp;amp; evolve using &amp;quot;branch&amp;quot; method&lt;br /&gt;
## Evolve 3D art: [http://endlessforms.com Endless Forms]&lt;br /&gt;
## CPPN-NEAT Algorithm: Compositional Pattern Producing Networks (CPPNs)-NeuroEvolution of Augmenting Topologies (NEAT); [https://www.ncbi.nlm.nih.gov/pubmed/20964537 PicBreeder paper]; or [http://campbellssite.com/papers/secretan_chi08.pdf a PDF version]&lt;br /&gt;
# [https://www.nature.com/articles/s41586-018-0102-6 NeuroEvolution (by DeepMind team)]&lt;br /&gt;
# Robotic snake (and other soft robots)&lt;br /&gt;
## A controller built with physics laws difficult (too many parameters)&lt;br /&gt;
## Simulation with evolutionary computing: &lt;br /&gt;
### (Genotype) A list of 13 commands (one for each segment; each being a neural net, with 25 inputs and one output of joint angles)&lt;br /&gt;
### (Phenotype) Fitness function: total displacement&lt;br /&gt;
### Algorithm: [https://medium.com/@devonfulcher3/the-map-elites-algorithm-finding-optimality-through-diversity-def6dcbc0f5b MAP-Elites]; [https://www.nature.com/articles/nature14422 Nature paper]&lt;br /&gt;
### Approach: training with simulated data&lt;br /&gt;
# Robotic Knightfish&lt;br /&gt;
## [https://www.youtube.com/watch?v=3XjgZbs0t2g Youtube demo]&lt;br /&gt;
## Algorithm: [https://en.wikipedia.org/wiki/CMA-ES CMA-ES (Covariance Matrix Adaptation Evolution Strategy)]&lt;br /&gt;
### Genotype: 15 variables (sinusoidal wave function or Fourier Series)&lt;br /&gt;
### Phenotype/Fitness: speed&lt;br /&gt;
## Implementation: DEAP&lt;br /&gt;
# Compositional Protein design (CPD)&lt;br /&gt;
## Genotype: side-chain configuration determined by [https://www.ncbi.nlm.nih.gov/pubmed/8464064 rotamer library]&lt;br /&gt;
## Phenotype/Fitness: Rosetta energy function &amp;amp; functional (e.g., binding affinity)&lt;br /&gt;
## Fitness landscape: each node is a protein structure, each edge represent connections/relatedness&lt;br /&gt;
# Simulation-based optimization: [http://simopt.org/ Problem Sets]&lt;br /&gt;
# Self-evolving software([http://geneticprogramming.com/ Genetic Programming])&lt;br /&gt;
&lt;br /&gt;
==Week 2. Toy Problem: OneMax Optimization==&lt;br /&gt;
# Problem: Create a list of L random bits (0 or 1). Evolve the list until the fitness reaches the maximum value (i.e., contains only 1&#039;s)&lt;br /&gt;
# Neutral Evolution:&lt;br /&gt;
## Create a vector of L random bits (e.g, L=20). Hint for creating a random vector of 0&#039;s and 1&#039;s in R: &amp;lt;code&amp;gt;ind&amp;lt;-sample(c(0,1), prob=c(0.5.0.5), replace=T, size=20)&amp;lt;/code&amp;gt;. Biologically, this vector represents a single haploid genome with 20 loci, each with two possible alleles (0 or 1).&lt;br /&gt;
## Create a population of N=100 such individuals. Hint: creating a list of vectors in R: &amp;lt;code&amp;gt;pop &amp;lt;- lapply(1:100, function(x){&amp;lt;insert sample() function above&amp;gt;})&amp;lt;/code&amp;gt;&lt;br /&gt;
## For each generation, each individual reproduces 10 gametes with mutation (with the probability for bit flip: mu = 1/L = 0.1). Hint: write a mutation function that takes an individual as input and outputs a mutated gamete. Use the &amp;quot;broken stick&amp;quot; algorithm to implement mutation rate: &amp;lt;code&amp;gt;cutoff &amp;lt;- runif(1); ifelse(cutoff &amp;lt;= mu, flip-bits, no-flip)&amp;lt;/code&amp;gt;&lt;br /&gt;
## Take a random sample of N=100 gametes into the next generation &amp;amp; repeat above&lt;br /&gt;
## Plot mean, minimum, and maximum fitness over generations&lt;br /&gt;
## Plot diversity statistics over generation (including average allelic heterozygosity per locus as well as haplotype heterozygosity). Hint: write two functions for these heterozygosity&lt;br /&gt;
# Add natural selection (proportional scheme)&lt;br /&gt;
## Individuals reproduce with fitness proportional to the total number of 1&#039;s.&lt;br /&gt;
## Iterate until a population contains at least one individual with all 1&#039;s&lt;br /&gt;
## Plot mean, minimum, and maximum fitness over generations&lt;br /&gt;
## Plot diversity (expectation: decreasing)&lt;br /&gt;
# Add natural selection (tournament scheme)&lt;br /&gt;
## Randomly selecting n=5 individuals and allow the fittest one to make gametes&lt;br /&gt;
## Plot mean, minimum, and maximum fitness over generations (Expectation: faster optimization)&lt;br /&gt;
## Plot diversity (expectation: decreasing fasters)&lt;br /&gt;
# Add crossover&lt;br /&gt;
## Hint: write a crossover function&lt;br /&gt;
## Does it reach optimization faster?&lt;br /&gt;
# Code submissions&lt;br /&gt;
## [https://github.com/JohnDi0505/MonteCarlo_Simulation-Biostats/blob/master/OneMax%20Optimization.ipynb Python Notebook by John]&lt;br /&gt;
## [http://rpubs.com/weigang/407979 R code by Weigang]&lt;br /&gt;
## [http://rpubs.com/ChrisNSP/408534 R code by Panlasigui]&lt;br /&gt;
## [https://github.com/Inhenn/Toy_Problem_Bio/blob/master/Bio_Toy.ipynb Python Notebook by Yinheng]&lt;br /&gt;
## [http://rpubs.com/desiree/409200 RPub by Desiree]&lt;br /&gt;
==Week 3. Python &amp;amp; R Packages for evolutionary computation==&lt;br /&gt;
* [http://deap.readthedocs.io/en/master/ DEAP: a Python package for Evolutionary Computing]&lt;br /&gt;
** Look under [http://deap.readthedocs.io/en/master/examples/index.html Examples] to repeat the OneMax code&lt;br /&gt;
* [https://cran.r-project.org/web/packages/GA/vignettes/GA.html GA: An R package for genetic programming]&lt;br /&gt;
** Reference 1 (Examples in Section 4). [https://www.jstatsoft.org/v53/i04/ Scrucca, L. (2013) GA: A Package for Genetic Algorithms in R.]&lt;br /&gt;
** Reference 2 (for Advanced applications). [https://journal.r-project.org/archive/2017/RJ-2017-008 Scrucca, L. (2017) On some extensions to GA package: hybrid optimisation, parallelisation and islands evolution. ]&lt;br /&gt;
** OneMax code &amp;amp; plots (Weigang)&lt;br /&gt;
** Example 4.1a. One variable optimization: f(x) = |x| + cos(x) (Muhammad and Desiree)&lt;br /&gt;
** Example 4.1b. One variable optimization: f(x) = (x2 + x) cos(x)&lt;br /&gt;
** Example 4.2.  Two-parameter optimization: f(x;y) = 20 + x^2 + y^2 * 10(cos(2x) + cos(2y)) (Muhammad and Desiree)&lt;br /&gt;
** Example 4.3.  Curve-fitting: tree growth&lt;br /&gt;
** Example 4.7.  Constrained optimization: Knapsack Problem&lt;br /&gt;
** Example 4.8.  Combinatorial optimization: Traveling Salesman Problem (Brian)&lt;br /&gt;
** Advanced application 1. Stock portfolio (Hybrid algorithms)&lt;br /&gt;
** Advanced application 2. Parallelization&lt;br /&gt;
** Advanced application 3. Island model&lt;br /&gt;
* Code submissions&lt;br /&gt;
** [http://rpubs.com/weigang/410426 rPub for OneMax (by Weigang)]&lt;br /&gt;
** One &amp;amp; Two-dimensional functional optimization with GA (by Desiree &amp;amp; Mohamud) : [http://rpubs.com/desireepante/411065 Entropy function]&lt;br /&gt;
** [https://github.com/Inhenn/Knapsack-Problem-Using-DEAP/blob/master/DEAP3.ipynb OneMax and Knapsack Problem with DEAP (by Yinheng)]&lt;br /&gt;
** Traveling Salesman Problem with GA (by Brian)&lt;br /&gt;
&lt;br /&gt;
==Week 4. Multiplex Problem==&lt;br /&gt;
==Week 5. Genetic Programming==&lt;br /&gt;
&lt;br /&gt;
=Season II. Summer 2017 (Theme: Machine Learning)=&lt;br /&gt;
==Week 1. Introduction &amp;amp; the backprop algorithm==&lt;br /&gt;
[[File:Iris-box3.png|thumbnail]] [[File:Iris-box4.png|thumbnail]]&lt;br /&gt;
# Problem: Classification/Clustering/Predication of flower species (a total of 3 possible species in the sample data set) based on four phenotypic traits/measurements&lt;br /&gt;
# The &amp;quot;iris&amp;quot; data set: exploratory data analysis with visualization &amp;amp; descriptive statistics: &amp;lt;code&amp;gt;data(&amp;quot;iris&amp;quot;); plot()&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;summary()&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;qqnorm(); qqline(); hist()&amp;lt;/code&amp;gt;; normalization with &amp;lt;code&amp;gt;scale()&amp;lt;/code&amp;gt; (Roy)&lt;br /&gt;
# Mathematics of backpropagating errors to neural connections (Oliver)&lt;br /&gt;
## Objective/optimization function measuring difference between target (&amp;lt;code&amp;gt;t&amp;lt;/code&amp;gt;, expected) and neural activity &amp;lt;code&amp;gt;y&amp;lt;/code&amp;gt;: &amp;lt;code&amp;gt;G=Sum{t*log(y)+(1-t)log(1-y)}&amp;lt;/code&amp;gt; (this is known as the &amp;quot;cross-entropy&amp;quot; error function; the other alternative is &amp;quot;minimal squared error (MSE)&amp;quot;), which has the gradient in the simple form of &amp;lt;code&amp;gt;g=-(t-y)x, where x is the input&amp;lt;/code&amp;gt;. The objective function is minimized when weights are updated by the gradient. Error-minimization by MSE has similar effects but harder to calculate.&lt;br /&gt;
## Learning algorithm is presented&lt;br /&gt;
&lt;br /&gt;
==Week 2. Traditional approaches to multivariate clustering/classification==&lt;br /&gt;
# Dimension reduction with Multidimensional Scaling &amp;lt;code&amp;gt;cmdscale()&amp;lt;/code&amp;gt; [http://rpubs.com/meibyderp/281616 Mei&#039;s rNoteBook]. &lt;br /&gt;
# Dimension reduction with Principal Component Analysis&amp;lt;code&amp;gt;princomp()&amp;lt;/code&amp;gt;. [http://rpubs.com/ssipa/281715 Sipa&#039;s rNoteBook]&lt;br /&gt;
# Multivariate clustering with Hierarchical Clustering &amp;lt;code&amp;gt;hclust()&amp;lt;/code&amp;gt; (Saymon)&lt;br /&gt;
# Multivariate clustering with k-means &amp;lt;code&amp;gt;kmeans()&amp;lt;/code&amp;gt; [http://rpubs.com/roynunez/281963 Roy&#039;s rNoteBook]&lt;br /&gt;
# Classification based on logistic regression &amp;lt;code&amp;gt;glm()&amp;lt;/code&amp;gt;, linear discriminatory analysis &amp;lt;code&amp;gt;lda()&amp;lt;/code&amp;gt;, &amp;amp; k-nearest neighbor &amp;lt;code&amp;gt;library(class); knn()&amp;lt;/code&amp;gt; (John)&lt;br /&gt;
# Modern, non-linear classifiers: Decision Trees (DT), Support Vector Machines (SVM), and Artificial Neural Networks (ANN) (Brian)&lt;br /&gt;
&lt;br /&gt;
==Week 3. Single-neuron classifier==&lt;br /&gt;
# Algorithm: [[File:Ml-image-1a.jpg|thumbnail]]&lt;br /&gt;
## Read input data with &amp;lt;code&amp;gt;N=150 flowers&amp;lt;/code&amp;gt;. Reduce to a matrix &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; with two traits (use trait 1 &amp;amp; 3) and two species (use rows 51-150, the last two species, skip the first species [easy to separate]) for simplicity. Create a target vector &amp;lt;code&amp;gt;t &amp;lt;- c(rep(0,50),rep(1,50))&amp;lt;/code&amp;gt; indicating two species&lt;br /&gt;
## Initialize the neuron with two random weights &amp;lt;code&amp;gt;w&amp;lt;-runif(2, 1e-3, 1e-2)&amp;lt;/code&amp;gt; and one random bias &amp;lt;code&amp;gt;b&amp;lt;-runif(1)&amp;lt;/code&amp;gt;&lt;br /&gt;
## Neuron activation with &amp;lt;code&amp;gt;k=2&amp;lt;/code&amp;gt; connection weights: &amp;lt;code&amp;gt;a=sum(x[k] * w[k])&amp;lt;/code&amp;gt;&lt;br /&gt;
## Neuron activity/output: &amp;lt;code&amp;gt;y=1/(1+exp(-a-b))&amp;lt;/code&amp;gt; (This logistic function ensures output values are between zero and one)&lt;br /&gt;
## Learning rules: learning rate &amp;lt;code&amp;gt;eta=0.1&amp;lt;/code&amp;gt;, backpropagate error (&amp;quot;e&amp;quot;) to get two updated weights (for individual &amp;lt;code&amp;gt;i&amp;lt;/code&amp;gt;, feature &amp;lt;code&amp;gt;k&amp;lt;/code&amp;gt;): &amp;lt;code&amp;gt;e[i]=t[i]-y[i]; g[k,i]= -e[i] * x[k,i]; g.bias[i] = -e[i] for bias&amp;lt;/code&amp;gt;; Batch update weights &amp;amp; bias: &amp;lt;code&amp;gt;w[k]=w[k] - eta * sum(g[k,i]); b = b - eta * sum(g.bias[i])&amp;lt;/code&amp;gt; (same rule for the bias parameter &amp;lt;code&amp;gt;b&amp;lt;/code&amp;gt;); Repeat/update for &amp;lt;code&amp;gt;L=1000 epochs (or generations)&amp;lt;/code&amp;gt;&lt;br /&gt;
## Output: weights and errors for each epoch&lt;br /&gt;
# Evaluation:&lt;br /&gt;
## Plot changes of weights over epoch&lt;br /&gt;
## Use the last weights and bias to predict species&lt;br /&gt;
## Compare prediction with target to get accuracy&lt;br /&gt;
## Plot scatter plot (x2 vs x1) and add a line using the weights &amp;amp; bias at epoch=1,50,100,200,500, 1000: &amp;lt;code&amp;gt;a=w1*x1 + w2*x2 + b; a=0&amp;lt;/code&amp;gt;. The lines should show increasing separation of the two species.&lt;br /&gt;
# Code submissions&lt;br /&gt;
## Roy: [http://rpubs.com/roynunez/283699 rPubs notebook]&lt;br /&gt;
## John: [https://github.com/JohnDi0505/MonteCarlo_Simulation-Biostats/blob/master/Neuron_Networks/Single_Neuron_Classification.py Python code on github]&lt;br /&gt;
## Mei: [http://rpubs.com/meibyderp/284498 rPubs notebook]&lt;br /&gt;
## Brian: [http://rpubs.com/drtwisto/286938 rPubs notebook]&lt;br /&gt;
## Weigang: [http://rpubs.com/weigang/279898 rPubs notebook]&lt;br /&gt;
# Questions for future exploration&lt;br /&gt;
## How to avoid over-fitting by regularization (a way to penalize model complexity by adding a weight decay parameter &amp;lt;code&amp;gt;alpha&amp;lt;/code&amp;gt;)&lt;br /&gt;
## Bayesian confidence interval of predictions (with Monte Carlo simulation)&lt;br /&gt;
## Limitations: equivalent to PCA (linear combination of features); adding a hidden layer generalize the neural net to be a non-linear classifier&lt;br /&gt;
&lt;br /&gt;
==Week 4. Single-layer, multiple-neuron learner==&lt;br /&gt;
[[File:Multiple-neuron-learner.png|thumbnail]]&lt;br /&gt;
# Predict all three species (with &amp;quot;one-hot&amp;quot; coding) using all four features: e.g., 100 for species 1, 010 for species 2, and 001 for species 3. Create 3 neurons, each one outputting one digit.&lt;br /&gt;
# Use three neurons, each accepts 4 inputs and output 1 activity&lt;br /&gt;
# Use softmax to normalize the final three output activities&lt;br /&gt;
# Code submissions&lt;br /&gt;
## John: &lt;br /&gt;
### [http://www.kdnuggets.com/2016/07/softmax-regression-related-logistic-regression.html reference this webpage]&lt;br /&gt;
### [https://github.com/JohnDi0505/MonteCarlo_Simulation-Biostats/blob/master/Neuron_Networks/Multiple_Neuron_Classification.py Python code]&lt;br /&gt;
### TensorFlow code&lt;br /&gt;
## Roy: [http://rpubs.com/roynunez/288023 rpubs notebook]&lt;br /&gt;
## Mei: [http://rpubs.com/meibyderp/288400 rPubs Notebook]&lt;br /&gt;
## Weigang: [http://rpubs.com/weigang/287025 rPubs notebook]&lt;br /&gt;
# Challenge: Implementation with TensorFlow, an Open Source python machine learning library by Google Deep Mind team. Follow [https://www.tensorflow.org/get_started/mnist/beginners this example of softmax regression]&lt;br /&gt;
# A biological application for the summer: identify SNPs associated with biofilm/swarming behavior in Pseudomonas&lt;br /&gt;
&lt;br /&gt;
==Week 5. Multi-layer (&amp;quot;Deep&amp;quot;) Neuron Network==&lt;br /&gt;
[[File:Iris-nnet.png|thumbnail]]&lt;br /&gt;
[[File:Nnet-overfitting.png|thumbnail|An investigation of under-fitting (at N=1) &amp;amp; over-fitting (at N=3 nodes). N=0 implies linear fitting.]]&lt;br /&gt;
# Input layer: 4 nodes (one for each feature)&lt;br /&gt;
# Output layer: 3 nodes (one for each species)&lt;br /&gt;
# Hidden layer: 2 hidden nodes&lt;br /&gt;
# Advantage: allows non-linear classification; using hidden layers (&amp;quot;convoluted neural net&amp;quot;) is able to capture high-order patterns&lt;br /&gt;
# Implementation: I&#039;m not going to hard-code the algorithm from scratch. The figure was produced using the R package with the following code: &amp;lt;code&amp;gt;library(nnet); library(neuralnet); targets.nn &amp;lt;- class.ind(c(rep(&amp;quot;setosa&amp;quot;,50), rep(&amp;quot;versicolor&amp;quot;,50), rep(&amp;quot;virginica&amp;quot;,50))) # 1-of-N encoding; iris.net &amp;lt;- neuralnet(formula = setosa + versicolor + virginica ~ Sepal.Length + Sepal.Width + Petal.Length + Petal.Width, data = training.iris, hidden = 2, threshold = 0.01, linear.output = T); plot(iris.net, rep=&amp;quot;best&amp;quot;)&amp;lt;/code&amp;gt;. It achieved 98.0% accuracy.&lt;br /&gt;
# Deep neural net allows non-linear, better fitting, but we don&#039;t want over-fitting by adding more hidden layers (or more neurons in the hidden layer). Identify under- and over-fitting with the following procedure:&lt;br /&gt;
## Randomly sample 100 as training set and the remaining as target. Repeat 100 times&lt;br /&gt;
## For each sample, plot accuracy for the training set, as well as accuracy for the target&lt;br /&gt;
## Find the point with the right balance of under- and over-fitting&lt;br /&gt;
&lt;br /&gt;
==Week 6. Unsupervised Neural Learner: Hopfield Networks==&lt;br /&gt;
# Hebbian model of memory formation (MacKay, Chapter 42). Preparatory work:&lt;br /&gt;
## Construct four memories, each for a letter (&amp;quot;D&amp;quot;, &amp;quot;J&amp;quot;, &amp;quot;C&amp;quot;, and &amp;quot;M&amp;quot;) using a 5-by-5 grid, with &amp;quot;-1&amp;quot; indicating blank space, and &amp;quot;1&amp;quot; indicating a pixel. Flatten the 5-by-5 matrix to a one-dimensional 1-by-25 vector (tensor). &lt;br /&gt;
## Write two functions, one to show a letter &amp;lt;code&amp;gt;show.letter(letter.vector)&amp;lt;/code&amp;gt;, which draws a pixel art of letters (print a blank if if -1, a &amp;quot;x&amp;quot; if 1), another to mutate the letter &amp;lt;code&amp;gt;mutate(letter.vector, number.pixel.flips)&amp;lt;/code&amp;gt;&lt;br /&gt;
# [https://en.wikipedia.org/wiki/Hopfield_network Hopfield Network]&lt;br /&gt;
## Store the four memories into a weight matrix, which consists of symmetric weights between neurons i and neuron j, i.e., w[i,j] = w[j,i]. We will use a total of 25 neurons, one for each pixel.&lt;br /&gt;
## First, combine the four vectors (one for each letter) into a matrix: &amp;lt;code&amp;gt;x &amp;lt;- matrix(c(letter.d, letter.j, letter.c, letter.m), nrow = 4, byrow = T)&amp;lt;/code&amp;gt;&lt;br /&gt;
## Second, calculate weight matrix by obtaining the outer product of a matrix multiplication: &amp;lt;code&amp;gt;w &amp;lt;- t(x) %*% x&amp;lt;/code&amp;gt;. Set diagonal values to be all 0&#039;s (i.e., to remove all self-connections): &amp;lt;code&amp;gt;for (i in 1:25) { w[i,i] = 0 }&amp;lt;/code&amp;gt;. This implements [https://en.wikipedia.org/wiki/Hebbian_theory Hebb learning], which translates correlations into strength of connections quantified by weights: large positive weights indicate mutual stimulation (e.g., 1 * 1 = 1 [to wire/strengthen the connection of co-firing neurons, and ...], -1 * -1 = 1 [to wire/strengthen the connection of co-inhibitory neurons as well]), large negative weights indicate mutual inhibition (e.g., 1 * -1 = -1 [to unwire/disconnect oppositely activated neurons]), and small weights indicate a weak connection (e.g., 0 * 1 = 0 [to weaken connections between neurons with uncorrelated activities, but do not unwire them]). (0, 1, and -1 being values of neuron activities)&lt;br /&gt;
## Third, implement the learning rule by updating activity for each neuron, sequentially (asynchronously): &amp;lt;code&amp;gt;a[i] &amp;lt;- sum(w[i,j] * x[j]) &amp;lt;/code&amp;gt;&lt;br /&gt;
## Iterate the previous step 1-5 times, your code should be able to (magically) restore the correct letter image even when the letter is mutated by 1-5 mutations in pixels. This exercise simulates the error-correction ability of memory (e.g., self-correcting encoding in CDs, a neon-light sign missing a stroke, or our ability to read/understand/reconstruct texts with typos, e.g., you have no problem reading/understanding this sentence: &amp;quot;It deosn’t mttaer in what order the ltteers in a word are, the olny iprmoatnt thing is taht the frist and lsat ltteer be in the rghit pclae.&amp;quot;). &lt;br /&gt;
## Expected capacity of a Hopfield Network: &amp;lt;code&amp;gt;number_of_memories = 0.138 * number_of_neurons&amp;lt;/code&amp;gt;. In our case, 25 (neurons) *0.138 = 3.45 memorized letters. So the network/memory fails if we squeeze in one additional letter (try it!).&lt;br /&gt;
# Code submissions&lt;br /&gt;
## Roy: [http://rpubs.com/roynunez/288433 rPubs Notebook]&lt;br /&gt;
## John: [https://github.com/JohnDi0505/MonteCarlo_Simulation-Biostats/blob/master/Neuron_Networks/Unsupervised_ML_Hopfield_Networks.ipynb Python Code on Github]&lt;br /&gt;
## Weigang: [http://rpubs.com/weigang/281806 rPubs Notebook]&lt;br /&gt;
# Potential biological applications: genetic code optimization; gene family identification&lt;br /&gt;
&lt;br /&gt;
==Week 7. Biological Applications==&lt;br /&gt;
A conceptual map for choosing ML algorithms: [[File:Ml map.png|thumbnail|Conceptual map from SciKit-Learn website]]&lt;br /&gt;
* The Python SciKit Learn framework may be tool of choice (instead of R). [http://scikit-learn.org/stable/auto_examples/index.html See these nice examples]  &lt;br /&gt;
* Bayesian Network (using e.g., R Package &amp;lt;code&amp;gt;bnlearn&amp;lt;/code&amp;gt;) to identify cell-signaling networks. Sache et al (2005). [http://science.sciencemag.org/content/308/5721/523.long Causal Protein-Signaling Networks Derived from Multiparameter Single-Cell Data]&lt;br /&gt;
* Identify proteins associated with learning in mice (Classification &amp;amp; Clustering)[https://archive.ics.uci.edu/ml/datasets/Mice+Protein+Expression UCI Mice Protein Expression Dataset]&lt;br /&gt;
* Predict HIV-1 protease cleavage sites (Classification): [https://archive.ics.uci.edu/ml/datasets/HIV-1+protease+cleavage UCI HIV-1 Protease Dataset]&lt;br /&gt;
* Lab project 1. Identify SNPs in 50 c-di-GMP pathway genes that are associated with c-di-GMP levels, swarming ability, and biofilm ability in 30 clinical isolates of Pseudomonas aeruginosa. Approach: supervised neural network (with regularization)&lt;br /&gt;
* Lab project 2. Identify genetic changes contributing to antibiotic resistance in 3 cancer patients. Approach: whole-genome sequencing followed by statistical analysis&lt;br /&gt;
* Lab project 3. Simulated evolution of genetic code. Approach: Multinomial optimization with unsupervised neural network&lt;br /&gt;
** An implementation example: [http://www.hoonzis.com/neural-networks-f-xor-classifier-and/ in C# language]&lt;br /&gt;
==Summer Project 1. Systems evolution of biofilm/swarming pathway (with Dr Joao Xavier of MSKCC)==&lt;br /&gt;
[[File:sim-cor-2.png|thumbnail| &amp;lt;b&amp;gt;Fig.1 .Simulated CDG-correlated SNPs.&amp;lt;/b&amp;gt; t-test results: (1) cor=0.8 (strong), t=-5.1142, df=95.596, p=1.619e-06; (2) cor=0.5 (medium), t=-4.7543, df=85.796, p=7.953e-06; (3) cor=0.2 (weak), t=-0.94585, df=79.28, p= 0.3471.]]&lt;br /&gt;
[[File:Sim-tree-snp-1.png|thumbnail|&amp;lt;b&amp;gt;Fig.2. Simulated tree-based SNPs.&amp;lt;/b&amp;gt; Generated with the APE function: &amp;lt;code&amp;gt;replicate(10, rTraitDisc(tr, states = c(0,1), rate = 100, model = &amp;quot;ER&amp;quot;))&amp;lt;/code&amp;gt;]]&lt;br /&gt;
# Acknowledgement: NSF Award 1517002&lt;br /&gt;
# Explanatory variable (genotypes): Whole-genome data (of ~30 clinical isolates &amp;amp; many experimentally evolved strains) as independent variables&lt;br /&gt;
# Explanatory variable (genotypes): ~50 genes related to cyclic-di-GMP synthesis and regulation (~8000 SNPs, ~1700 unique, ~5-8 major groups)&lt;br /&gt;
# Response variables (phenotypes): biofilm size, swarming size, antibiotic sensitivity profile, metabolomics measurements, c-di-GMP levels&lt;br /&gt;
# Sub-project 1: Database updates (Usmaan, Edgar, Christopher; continuing the work of Rayees and Raymond in previous years)&lt;br /&gt;
## c-di-GMP pathway SNPs in the new table &amp;quot;cdg_snp&amp;quot;&lt;br /&gt;
## c-di-GMP levels in &amp;quot;phenotype&amp;quot; table&lt;br /&gt;
## capture fig orthologs (in progress)&lt;br /&gt;
## capture matebolomics data with a new table (in process)&lt;br /&gt;
### [http://diverge.hunter.cuny.edu/~weigang/hm_metabolite-dendoclust.html a heatmap of deviation from mean (among strains of the same metabolite)]&lt;br /&gt;
### [http://diverge.hunter.cuny.edu/~weigang/matolites-fc-volcano-plot.html a volcano plot (p values based on t-test from group mean among strains of the same metabolite)]&lt;br /&gt;
# Sub-project 2. Supervised learning for predicting genes, SNPs associated with phenotypes&lt;br /&gt;
## Advantages over traditional regression analysis: ability to discover non-linear correction structure&lt;br /&gt;
## Challenges:&lt;br /&gt;
### over-fitting (we have only ~30 observations while thousands of SNPs, &amp;quot;curse of dimensionality&amp;quot;)&lt;br /&gt;
### the &amp;quot;effective&amp;quot; sample size is further discounted/reduced by phylogenetic relatedness among the strains.&lt;br /&gt;
## Single neuron, two-levels; by SNP groups (using &amp;lt;code&amp;gt;hclust()&amp;lt;/code&amp;gt;); by PCA (linear combination of SNPs); or by linkage groups (looking for homoplasy) (Mei), or by t-SNE (as suggested by Rayees)&lt;br /&gt;
## Three neurons, three-levels; by gene (Roy)[http://rpubs.com/roynunez/291340 july-14_nnet_cidigmp.R]&lt;br /&gt;
## To Do: (1) predict continuous-value targets; (2) add hidden layers; (3) correct for phylogenetic auto-correlation&lt;br /&gt;
# Sub-project 3. Generate simulated data (with Choleski Decomposition) (Weigang; continue the work in previous year by Ishmere, Zwar, &amp;amp; Rayees)&lt;br /&gt;
Fig.1. Simulated SNPs associated with cdg levels (w/o tree) [https://stats.stackexchange.com/questions/12857/generate-random-correlated-data-between-a-binary-and-a-continuous-variable inspired by this algorithm]&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
# simulate SNP-cdg correlation&lt;br /&gt;
r &amp;lt;- 0.2 # desired correlation coefficient&lt;br /&gt;
sigma &amp;lt;- matrix(c(1,r,r,1), ncol=2) # var-covariance matrix&lt;br /&gt;
s &amp;lt;- chol(sigma) # choleski decomposition&lt;br /&gt;
n &amp;lt;- 100 # number of random deviates (data points)&lt;br /&gt;
z &amp;lt;- s %*% matrix(rnorm(n*2), nrow=2) # 100 correlated normally distributed deviates with cor(x,y)=r&lt;br /&gt;
u &amp;lt;- pnorm(z) # get probabilities for each deviates&lt;br /&gt;
snp.states &amp;lt;- qbinom(u[1,], 1, 0.5) # discretize the 1st vector of probabilities into 0/1 with Bernoulli trial&lt;br /&gt;
idx.0 &amp;lt;- which(snp.states == 0); # indices for &amp;quot;0&amp;quot;&lt;br /&gt;
idx.1 &amp;lt;- which(snp.states == 1); # indices for &amp;quot;1&amp;quot;&lt;br /&gt;
# boxplots with stripcharts&lt;br /&gt;
boxplot(u[2,] ~ snp.states, main=&amp;quot;cor=0.2&amp;quot;, xlab=&amp;quot;SNP states&amp;quot;, ylab=&amp;quot;CDG level&amp;quot;)&lt;br /&gt;
stripchart(u[2,] ~ snp.states, vertical=T, pch=1, method=&amp;quot;jitter&amp;quot;, col=2,  add=T)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
Fig.2. Simulated SNPs associated with strain phylogeny&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
# Simulate SNPs on a tree&lt;br /&gt;
tr &amp;lt;- read.tree(&amp;quot;cdg-tree-mid.dnd&amp;quot;) # mid-point rooted tree&lt;br /&gt;
X &amp;lt;- replicate(10, rTraitDisc(tr, states = c(0,1), rate = 100, model = &amp;quot;ER&amp;quot;))&lt;br /&gt;
id &amp;lt;- read.table(&amp;quot;cdg.strains.txt3&amp;quot;, row.names = 1, sep=&amp;quot;\t&amp;quot;)&lt;br /&gt;
par.tr &amp;lt;- plot(tr, no.margin = T, x.lim = 0.03, show.tip.label = F)&lt;br /&gt;
text(rep(0.013,30), 1:30, id[tr$tip.label,1], pos = 4, cex=0.75)&lt;br /&gt;
text(rep(0.02,30), 1:30, snps[tr$tip.label], pos=4, cex=0.75)&lt;br /&gt;
add.scale.bar()&lt;br /&gt;
abline(h=1:30, col=&amp;quot;gray&amp;quot;, lty=2)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
To Do/Challenges: (1) simulate multiple SNPs (linear correlation); (2) simulate epistasis (non-linear correlation); (3) simulate phylogenetic correlation&lt;br /&gt;
* Simulation &amp;amp; xgboost code contributed by Mei &amp;amp; Yinheng (January 2018)&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
&lt;br /&gt;
#To create a simulated matrix w/ correlated xy variables; NOTE: ONLY 1 correlated x variable&lt;br /&gt;
get1Simulated &amp;lt;- function(corco, nstrains, snps){&lt;br /&gt;
  r &amp;lt;- corco/10 # desired correlation coefficient&lt;br /&gt;
  sigma &amp;lt;- matrix(c(1,r,r,1), ncol=2) # var-covariance matrix&lt;br /&gt;
  s &amp;lt;- chol(sigma) #cholesky decomposition&lt;br /&gt;
  n &amp;lt;- nstrains&lt;br /&gt;
  z &amp;lt;- s %*% matrix(rnorm(n*2), nrow=2)&lt;br /&gt;
  u &amp;lt;- pnorm(z[2,])&lt;br /&gt;
  snp.states &amp;lt;- qbinom(u, 1, 0.5)&lt;br /&gt;
&lt;br /&gt;
  known &amp;lt;- t(rbind(z[1,],snp.states))&lt;br /&gt;
  rand &amp;lt;- pnorm(matrix(rnorm(nstrains*(snps-1)),nrow = nstrains))&lt;br /&gt;
  snp.rand&amp;lt;-qbinom(rand,1,0.5)&lt;br /&gt;
  x &amp;lt;- cbind(known, snp.rand)&lt;br /&gt;
&lt;br /&gt;
  colnames(x)[1] &amp;lt;- &amp;quot;target&amp;quot;&lt;br /&gt;
  colnames(x)[-1] &amp;lt;- paste(&amp;quot;feature&amp;quot;, seq_len(ncol(x)-1), sep = &amp;quot;.&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
  return(x)&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
#Use this function to artificially create 2 x variables that have relationship w/ Y.&lt;br /&gt;
getSimulated.chol &amp;lt;- function(snp1, snp2, nstrains, snps){&lt;br /&gt;
  n &amp;lt;- nstrains&lt;br /&gt;
&lt;br /&gt;
  btwn &amp;lt;- snp1 * snp2 #this is a suggested value for inter-SNP correlation. the covariance matrix has to be semi-positive infinite in order to carry out the decomposition&lt;br /&gt;
&lt;br /&gt;
  covar.m &amp;lt;- matrix(c(1.0, snp1, snp2, snp1, 1.0, btwn, snp2, btwn, 1.0), nrow = 3) #cor-matrix&lt;br /&gt;
&lt;br /&gt;
  s &amp;lt;- chol(covar.m)&lt;br /&gt;
  z &amp;lt;- s %*% matrix(rnorm(n*3), nrow = 3)&lt;br /&gt;
&lt;br /&gt;
  #decretize the 2 variables&lt;br /&gt;
  u &amp;lt;- t(pnorm(z[2:nrow(z),]))&lt;br /&gt;
  decretize.u &amp;lt;-qbinom(u,1,0.5)&lt;br /&gt;
&lt;br /&gt;
  #generate some random discrete variables and combine w/ the 2 correlated variables&lt;br /&gt;
  rand &amp;lt;- matrix(rnorm(nstrains*(snps-2)),nrow = nstrains)&lt;br /&gt;
  rand.u &amp;lt;- pnorm(rand)&lt;br /&gt;
  snp.rand &amp;lt;- qbinom(rand.u,1,0.5)&lt;br /&gt;
  x &amp;lt;- cbind(z[1,], decretize.u, snp.rand)&lt;br /&gt;
&lt;br /&gt;
  dimnames(x) &amp;lt;- list(c(), c(&amp;quot;target&amp;quot;, paste(&amp;quot;feature&amp;quot;, seq_len(ncol(x)-1), sep = &amp;quot;.&amp;quot;)))&lt;br /&gt;
&lt;br /&gt;
  return(x)&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
#This utilizes xgboost function w/o cross-validating the strains and w/ tuned parameters (we think it&#039;s the optimal set) tested by Yinheng&#039;s python [https://github.com/weigangq/mic-boost/blob/master/micboost/__init__.py getBestParameters] function. Please make sure to have xgboost installed. &lt;br /&gt;
runXg &amp;lt;- function(x){&lt;br /&gt;
  require(xgboost)&lt;br /&gt;
&lt;br /&gt;
  bst &amp;lt;- xgboost(data = x[,2:ncol(x)], label = x[,1], max.depth = 2, eta = .05, gamma = 0.3, nthread = 2, nround = 10, verbose = 0, eval_metric = &amp;quot;rmse&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
  importance_matrix &amp;lt;- xgb.importance(model = bst, feature_names = colnames(x[,2:ncol(x)]))&lt;br /&gt;
&lt;br /&gt;
  return(importance_matrix)&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
#Same as above but w/ cross validation.&lt;br /&gt;
runXg.cv &amp;lt;- function(x){&lt;br /&gt;
  require(caret)&lt;br /&gt;
  require(xgboost)&lt;br /&gt;
&lt;br /&gt;
  ind &amp;lt;- createDataPartition(x[,1], p = 2/3, list = FALSE )&lt;br /&gt;
&lt;br /&gt;
  trainDf &amp;lt;- as.matrix(x[ind,])&lt;br /&gt;
  testDf &amp;lt;- as.matrix(x[-ind,])&lt;br /&gt;
&lt;br /&gt;
  dtrain &amp;lt;- xgb.DMatrix(data = trainDf[,2:ncol(trainDf)], label = trainDf[,1])&lt;br /&gt;
  dtest &amp;lt;- xgb.DMatrix(data = testDf[,2:ncol(testDf)], label = testDf[,1])&lt;br /&gt;
&lt;br /&gt;
  watchlist &amp;lt;- list(train=dtrain, test=dtest)&lt;br /&gt;
&lt;br /&gt;
  bst &amp;lt;- xgb.train(data=dtrain, nthread = 2, nround=10, watchlist=watchlist, eval.metric = &amp;quot;rmse&amp;quot;, verbose = 0)&lt;br /&gt;
&lt;br /&gt;
  importance_matrix &amp;lt;- xgb.importance(model = bst, feature_names = colnames(x[,2:ncol(x)]))&lt;br /&gt;
&lt;br /&gt;
  return(importance_matrix)&lt;br /&gt;
&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
#This function returns a master table with different iterated values that finds the best number of predictor variables (x).&lt;br /&gt;
getTable &amp;lt;- function(x, importance_matrix, cor){&lt;br /&gt;
  # v.names &amp;lt;- c(&amp;quot;feature&amp;quot;, &amp;quot;SNPs&amp;quot;, &amp;quot;strains&amp;quot;, &amp;quot;cor_given&amp;quot;, &amp;quot;cor_actual&amp;quot;, &amp;quot;gain&amp;quot;, &amp;quot;cover&amp;quot;, &amp;quot;rank&amp;quot;)&lt;br /&gt;
  v.names &amp;lt;- c(&amp;quot;feature&amp;quot;, &amp;quot;SNPs&amp;quot;, &amp;quot;strains&amp;quot;, &amp;quot;cor_given&amp;quot;, &amp;quot;cor_actual&amp;quot;, &amp;quot;p.value&amp;quot;, &amp;quot;gain&amp;quot;, &amp;quot;cover&amp;quot;, &amp;quot;rank&amp;quot;)&lt;br /&gt;
  for (i in 1:length(v.names)){&lt;br /&gt;
    assign(v.names[i], numeric())&lt;br /&gt;
  }&lt;br /&gt;
  featNum &amp;lt;- NULL&lt;br /&gt;
&lt;br /&gt;
  n &amp;lt;- 0&lt;br /&gt;
  new_cor &amp;lt;- rep(cor, length(x)/length(cor))&lt;br /&gt;
  for (num in 1:length(x)) {&lt;br /&gt;
    featNum &amp;lt;- NULL&lt;br /&gt;
    imp_matrix &amp;lt;- importance_matrix[[num]]&lt;br /&gt;
    x.iter &amp;lt;- x[[num]]&lt;br /&gt;
    x.pv &amp;lt;- getPV(x[[num]])&lt;br /&gt;
&lt;br /&gt;
    for (i in 1:2){&lt;br /&gt;
      featNum &amp;lt;- c(featNum, grep(paste(&amp;quot;feature.&amp;quot;,i,&amp;quot;$&amp;quot;, sep = &amp;quot;&amp;quot;), imp_matrix$Feature, perl = TRUE))&lt;br /&gt;
    }&lt;br /&gt;
&lt;br /&gt;
    n &amp;lt;-  n + 1&lt;br /&gt;
&lt;br /&gt;
    for(i in 1:2){&lt;br /&gt;
      f &amp;lt;- featNum[i]&lt;br /&gt;
      feature&amp;lt;-c(feature, imp_matrix$Feature[f])&lt;br /&gt;
      SNPs &amp;lt;-c(SNPs, 100)&lt;br /&gt;
      strains &amp;lt;- c(strains, nrow(x.iter))&lt;br /&gt;
      cor_given &amp;lt;- c(cor_given, new_cor[n])&lt;br /&gt;
      cor_actual &amp;lt;-c(cor_actual, cor(x.iter[,1],x.iter[,grep(paste(imp_matrix$Feature[f], &amp;quot;$&amp;quot;, sep = &amp;quot;&amp;quot;), colnames(x.iter), perl = TRUE)]))&lt;br /&gt;
      p.value &amp;lt;- c(p.value, x.pv[,imp_matrix$Feature[f]])&lt;br /&gt;
      gain &amp;lt;- c(gain, imp_matrix$Gain[f])&lt;br /&gt;
      cover &amp;lt;- c(cover, imp_matrix$Cover[f])&lt;br /&gt;
      rank &amp;lt;- c(rank, f)&lt;br /&gt;
    }&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
  x.master &amp;lt;- data.frame(feature, SNPs, strains, cor_given, cor_actual, p.value, gain, cover, rank)&lt;br /&gt;
&lt;br /&gt;
  return(x.master)&lt;br /&gt;
}&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Summer Project 2. Whole-genome variants associated with multidrug resistance in clinical Pseudomonas &amp;amp; E.coli isolates (with Dr Yi-Wei Tang of MSKCC)==&lt;br /&gt;
# Acknowledgement: Hunter CTBR Pilot Award&lt;br /&gt;
# Stage 1: Select patients &amp;amp; strains for whole-genome sequencing by MiSeq, based on drug sensitivities (16 isolates from 5 patients were isolated, tested, and selected by April, 2017)&lt;br /&gt;
# Stage 2: Genome sequencing (FASTQ files generated, 3 replicates for each isolates; done by June 2017)&lt;br /&gt;
# Stage 3: Variant call&lt;br /&gt;
## Reference strains identified using Kraken (Roy)&lt;br /&gt;
## VCF generation using cortex_var (Michele and Hanna, led by John)&lt;br /&gt;
# Stage 4: Variant annotation&lt;br /&gt;
# Stage 5: Statistical analysis&lt;br /&gt;
# Stage 6: Web report&lt;br /&gt;
&lt;br /&gt;
==Summer Project 3. Origin of Genetic Code==&lt;br /&gt;
Tools for testing SGC &amp;amp; evolved codes&lt;br /&gt;
* Shuffle code: &lt;br /&gt;
&amp;lt;syntaxhighlight lang=&#039;bash&#039;&amp;gt;&lt;br /&gt;
./shuffle-code.pl&lt;br /&gt;
    -f &amp;lt;sgc|code-file&amp;gt; (required)&lt;br /&gt;
    -s &amp;lt;1|2|3|4|5&amp;gt; (shuffle by 1st, 2nd, 3rd, aa blocks, and all random)&lt;br /&gt;
    -p(olarity; default)&lt;br /&gt;
    -h(ydropathy)&lt;br /&gt;
    -v(olume)&lt;br /&gt;
    -e(iso-electricity)&lt;br /&gt;
# Input: &#039;sgc&#039; (built-in) or an evolved code file consisting of 64 rows of &amp;quot;codon&amp;quot;-&amp;quot;position&amp;quot;&lt;br /&gt;
# Output: a code file consisting of 64 rows of &amp;quot;codon&amp;quot; - &amp;quot;aa&amp;quot; (to be fed into ./code-stats.pl)&lt;br /&gt;
# Usage examples:&lt;br /&gt;
./shuffle-code.pl -f &#039;sgc&#039; -s 5 # randomly permute SGC&lt;br /&gt;
./shuffle-code.pl -f &#039;evolved-code.txt&#039; -p # evolved code, AA assigned according to polarity (default)&lt;br /&gt;
./shuffle-code.pl -f &#039;evolved-code.txt&#039; -s 5 -h # evolved code, AA assigned according to hydrophobicity, shuffled randomly&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* Code statistics&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&#039;bash&#039;&amp;gt;&lt;br /&gt;
./code-stats.pl&lt;br /&gt;
        [-h]    (help)&lt;br /&gt;
        [-f &#039;sgc (default)|code_file&#039;]  (code)&lt;br /&gt;
        [-s &#039;pair (default)|fit|path&#039;]  (stats)&lt;br /&gt;
        [-p &#039;pol|hydro|vol|iso&#039;, default &#039;grantham&#039;]    (aa prop)&lt;br /&gt;
        [-i (ti/tv, default 5)]&lt;br /&gt;
        [-b: begin codon (&#039;TTT&#039;) -q: panelty for 1st (50); -w: panelty for 2nd (100)]   (options for path)&lt;br /&gt;
# Input: &#039;sgc&#039; (built-in) or a code file consisting of 64 rows of &amp;quot;codon&amp;quot; - &amp;quot;aa&amp;quot;&lt;br /&gt;
# Output: codon or code statistics&lt;br /&gt;
# Usage examples:&lt;br /&gt;
./code-stats.pl # all defaults: print single-mutation codon pairs, grantham distance, for SGC&lt;br /&gt;
./code-stats.pl -s &#039;fit&#039; -p &#039;pol&#039; # print code fitness according to polarity, for SGC (default)&lt;br /&gt;
./code-stats.pl -s &#039;path&#039; # print tour length for SGC&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;gallery perrow=&amp;quot;4&amp;quot;&amp;gt;&lt;br /&gt;
Sgc-polar.png|SGC: ordered by polarity&lt;br /&gt;
Sgc-hydro.png|SGC: ordered by hydropathy&lt;br /&gt;
Sgc-vol.png|SGC: ordered by volume&lt;br /&gt;
Sgc-iso.png|SGC: ordered by isoelectricity&lt;br /&gt;
&lt;br /&gt;
Path-code3.png|An evolved code (by Oliver)&lt;br /&gt;
code-john-2.png|An evolved code (by John). Round-trip (last position connected with the first) AA assigned according to polarity gradient &amp;amp; with &amp;quot;TAA&amp;quot; as the 1st.&lt;br /&gt;
code-brian-1.png|An evolved code (by Brian). One-way tour (last position NOT connected with the first). AA assigned according to polarity gradient &amp;amp; with &amp;quot;TAA&amp;quot; as the 1st.&lt;br /&gt;
code-error-v3.png|code robustness: &amp;lt;font color=&amp;quot;red&amp;quot;&amp;gt;SGC&amp;lt;/font&amp;gt;, &amp;lt;font color=&amp;quot;blue&amp;quot;&amp;gt;John&#039;s simulated code &amp;lt;/font&amp;gt;, &amp;lt;font color=&amp;quot;orange&amp;quot;&amp;gt;Brian&#039;s simulated code &amp;lt;/font&amp;gt;, histogram: 1000 random codes. X-axis: mean-squared error (standardized) caused by single-nt substitutions. Evolved codes perform better than SGC!&lt;br /&gt;
&lt;br /&gt;
Aa.pca.png|PC1 (mostly polarity and hydropathy, anti-correlated) &amp;amp; PC2 (the other two, positively correlated) of 4 AA metrics. PC1 and PC2 could be used for AA ranking as composite variables&lt;br /&gt;
tRNA-tree.png|A tRNA gene tree from Aeropyrum pernix (an Archaea). Sequences from [http://trna.bioinf.uni-leipzig.de/DataOutput/ an rRNA database]. Structual alignment available&lt;br /&gt;
tRNA-tree-2.png|tRNA gene tree for Pyrococcus horikoshii (another Archaea). Sequences from [http://gtrnadb.ucsc.edu/GtRNAdb2/index.html UCSD tRNA database]. (Not structurally aligned; only fasta seqs available)&lt;br /&gt;
tRNA-seq-dist.png|AA distance appears to be correlated with tRNA seq differences (especially at low distances)&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
# Participants: Oliver, John, Brian&lt;br /&gt;
# A representation mimicking TSP (see figure at right)&lt;br /&gt;
# Code to calculate total (mutational, ti/tv) path of SGC?&lt;br /&gt;
# Code to calculate energy of SGC?&lt;br /&gt;
# Randomize amino acid assignment and obtain distributions of path &amp;amp; energy&lt;br /&gt;
# Code to minimize total path length and energy?&lt;br /&gt;
# How to evolve a robust SGC? mutation bias + polarity/hydropathy + usage&lt;br /&gt;
&lt;br /&gt;
=Season I. Spring 2017 (Themes: Dueling Idiots/Digital Dice/Wright Fisher Process)=&lt;br /&gt;
==&amp;quot;Coalescence&amp;quot; (Backward simulation of Wright-Fisher process) (Due May 12, 2017)==&lt;br /&gt;
[[File:Coalescent-tree.png|thumbnail|([http://raven.iab.alaska.edu/~ntakebay/teaching/programming/coalsim/node1.html Source])]]&lt;br /&gt;
[[File:Coalescent-output-1.png|thumbnail]]&lt;br /&gt;
We will conclude Spring 2017 season with the coalescence simulation (in summer, we will start experimenting with simulation of systems evolution) &lt;br /&gt;
&lt;br /&gt;
Previously, we simulated Wright-Fisher process of genetic drift (constant pop size, no selection) starting from a founder population and end with the present population. This is not efficient because the program has to track each individual in each generation, although the majority of them do not contribute to the present population.&lt;br /&gt;
&lt;br /&gt;
Coalescence simulation takes the opposite approach of starting from the present sample of k individuals and trace backward in time to their most recent common ancestor (MRCA). Due to the nature of Poisson process, the waiting time from one coalescence event (at time T) to the next one (at time T-1) is exponentially distributed with a mean of (k choose 2) generations. This process is iterated until the last coalescent event.&lt;br /&gt;
&lt;br /&gt;
The classic text on coalescence is [http://home.uchicago.edu/rhudson1/popgen356/OxfordSurveysEvolBiol7_1-44.pdf Richard Hudson&#039;s chapter], which includes  (in Appendix) C codes for simulating tree and mutations. Hudson is also the author of &amp;lt;code&amp;gt;ms&amp;lt;/code&amp;gt;, the widely used coalescence simulator.&lt;br /&gt;
&lt;br /&gt;
Like the &amp;lt;code&amp;gt;ms 10 1 -T&amp;lt;/code&amp;gt; command, your code should output a tree of 10 individuals. [A lot tougher than I thought; can&#039;t figure it out in R; so I did it in Perl]. Similarly, you could use R, where the &amp;lt;code&amp;gt;ape&amp;lt;/code&amp;gt; package has a function called &amp;lt;code&amp;gt;rcoal()&amp;lt;/code&amp;gt; that generate a random coalescence tree. Try this command: &amp;lt;code&amp;gt;plot(rcoal(10))&amp;lt;/code&amp;gt; (first load library by running &amp;lt;code&amp;gt;library(ape)&amp;lt;/code&amp;gt;).&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
|&lt;br /&gt;
*By Weigang (First draft)&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;perl&amp;quot;&amp;gt;&lt;br /&gt;
#!/usr/bin/env perl&lt;br /&gt;
# Basic coalescence&lt;br /&gt;
# First, simulate coalescent events with recursion&lt;br /&gt;
# Second, use Bio::Tree to export a newick tree&lt;br /&gt;
use strict;&lt;br /&gt;
use warnings;&lt;br /&gt;
use Data::Dumper;&lt;br /&gt;
use Algorithm::Numerical::Sample  qw /sample/;&lt;br /&gt;
use Math::Random qw(random_exponential);&lt;br /&gt;
use Bio::Tree::Tree;&lt;br /&gt;
use Bio::Tree::Node;&lt;br /&gt;
&lt;br /&gt;
######################&lt;br /&gt;
# Initialize&lt;br /&gt;
######################&lt;br /&gt;
die &amp;quot;Usage: $0 &amp;lt;num-of-samples&amp;gt;\n&amp;quot; unless @ARGV == 1;&lt;br /&gt;
my $nsamp = shift @ARGV;&lt;br /&gt;
my @samples;&lt;br /&gt;
for (my $i=1; $i&amp;lt;=$nsamp; $i++) { push @samples, {id=&amp;gt;$i, parent=&amp;gt;undef, br=&amp;gt;0} }&lt;br /&gt;
my $ctr = $nsamp;&lt;br /&gt;
my $time = 0;&lt;br /&gt;
my @all_nodes;&lt;br /&gt;
&lt;br /&gt;
################################################################################&lt;br /&gt;
# Simulate events with exponential time intervals between two successive events&lt;br /&gt;
#################################################################################&lt;br /&gt;
&amp;amp;coal(\@samples, \$ctr, \$time);&lt;br /&gt;
#print Dumper(\@all_nodes);&lt;br /&gt;
#########################################&lt;br /&gt;
# Reconstitute into tree using Bio::Tree&lt;br /&gt;
##########################################&lt;br /&gt;
my %seen_node;&lt;br /&gt;
my $max_id=0;&lt;br /&gt;
my @nodes;&lt;br /&gt;
foreach (@all_nodes) {&lt;br /&gt;
    $max_id = ($_-&amp;gt;{parent} &amp;gt; $max_id) ? $_-&amp;gt;{parent} : $max_id;&lt;br /&gt;
    if ($seen_node{$_-&amp;gt;{id}}) { # child exist, previously as a parent, no branch length&lt;br /&gt;
        $seen_node{$_-&amp;gt;{id}}-&amp;gt;branch_length(sprintf &amp;quot;%.6f&amp;quot;, $_-&amp;gt;{br}); # add branch length&lt;br /&gt;
        if ($seen_node{$_-&amp;gt;{parent}}) { # parent exist&lt;br /&gt;
            $seen_node{$_-&amp;gt;{parent}}-&amp;gt;add_Descendent($seen_node{$_-&amp;gt;{id}});&lt;br /&gt;
        } else { # parent new&lt;br /&gt;
            my $pa = Bio::Tree::Node-&amp;gt;new(-id=&amp;gt;$_-&amp;gt;{parent});&lt;br /&gt;
            $pa-&amp;gt;add_Descendent($seen_node{$_-&amp;gt;{id}});&lt;br /&gt;
            $seen_node{$_-&amp;gt;{parent}} = $pa;&lt;br /&gt;
        }&lt;br /&gt;
    } else { # child new&lt;br /&gt;
        my $new = Bio::Tree::Node-&amp;gt;new(-id=&amp;gt;$_-&amp;gt;{id}, -branch_length =&amp;gt; sprintf &amp;quot;%.6f&amp;quot;, $_-&amp;gt;{br});&lt;br /&gt;
        $seen_node{$_-&amp;gt;{id}} = $new;&lt;br /&gt;
        if ($seen_node{$_-&amp;gt;{parent}}) { # parent exist&lt;br /&gt;
            $seen_node{$_-&amp;gt;{parent}}-&amp;gt;add_Descendent($new);&lt;br /&gt;
        } else { # parent new&lt;br /&gt;
            my $pa = Bio::Tree::Node-&amp;gt;new(-id=&amp;gt;$_-&amp;gt;{parent});&lt;br /&gt;
            $pa-&amp;gt;add_Descendent($new);&lt;br /&gt;
            $seen_node{$_-&amp;gt;{parent}} = $pa;&lt;br /&gt;
        }&lt;br /&gt;
    }&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
#my $root = $seen_node{$max_id};&lt;br /&gt;
my $tree=Bio::Tree::Tree-&amp;gt;new(-id=&amp;gt;&#039;coal-sim&#039;, -node=&amp;gt;$seen_node{$max_id}, -nodelete=&amp;gt;1);&lt;br /&gt;
print $tree-&amp;gt;as_text(&amp;quot;newick&amp;quot;), &amp;quot;\n&amp;quot;;&lt;br /&gt;
&lt;br /&gt;
exit;&lt;br /&gt;
&lt;br /&gt;
sub coal {&lt;br /&gt;
    my $ref_nodes = shift;&lt;br /&gt;
    my $ref_ct = shift;&lt;br /&gt;
    my $ref_time = shift;&lt;br /&gt;
    my $ct = $$ref_ct;&lt;br /&gt;
    my @current_nodes = @$ref_nodes;&lt;br /&gt;
    my $k = scalar @current_nodes;&lt;br /&gt;
    my @new_nodes;&lt;br /&gt;
    return unless $k &amp;gt; 1;&lt;br /&gt;
    my @pair = sample(-set =&amp;gt; $ref_nodes, -sample_size =&amp;gt; 2);&lt;br /&gt;
#    print $ct, &amp;quot;\t&amp;quot;, $$ref_time, &amp;quot;\t&amp;quot;, $pair[0]-&amp;gt;{id}, &amp;quot;\t&amp;quot;, $pair[1]-&amp;gt;{id}, &amp;quot;\n&amp;quot;;&lt;br /&gt;
    $$ref_time += random_exponential(1, 2/$k/($k-1));&lt;br /&gt;
    my $new_nd = {id=&amp;gt;$ct+1, parent=&amp;gt;undef, br=&amp;gt;$$ref_time};&lt;br /&gt;
    map {$_-&amp;gt;{parent} = $ct+1} @pair;&lt;br /&gt;
    map {$_-&amp;gt;{br} = $$ref_time - $_-&amp;gt;{br}} @pair;&lt;br /&gt;
    push @all_nodes, $_ for @pair;&lt;br /&gt;
    foreach (@current_nodes) {&lt;br /&gt;
        push @new_nodes, $_ unless $_-&amp;gt;{id} == $pair[0]-&amp;gt;{id} || $_-&amp;gt;{id} == $pair[1]-&amp;gt;{id};&lt;br /&gt;
    }&lt;br /&gt;
    push @new_nodes, $new_nd;&lt;br /&gt;
    $$ref_ct++;&lt;br /&gt;
    &amp;amp;coal(\@new_nodes, $ref_ct, $ref_time);&lt;br /&gt;
}&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
*By John&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;python&amp;quot;&amp;gt;&lt;br /&gt;
from Bio import Phylo&lt;br /&gt;
from io import StringIO&lt;br /&gt;
from random import sample&lt;br /&gt;
from scipy.misc import comb&lt;br /&gt;
from itertools import combinations&lt;br /&gt;
from numpy.random import exponential as exp&lt;br /&gt;
&lt;br /&gt;
# Generate Random Tree with Nodes &amp;amp; Waiting Time&lt;br /&gt;
individuals = [&amp;quot;A&amp;quot;, &amp;quot;B&amp;quot;, &amp;quot;C&amp;quot;, &amp;quot;D&amp;quot;, &amp;quot;E&amp;quot;, &amp;quot;F&amp;quot;, &amp;quot;G&amp;quot;, &amp;quot;H&amp;quot;, &amp;quot;I&amp;quot;, &amp;quot;J&amp;quot;]&lt;br /&gt;
node_dict = {}&lt;br /&gt;
reference = {}&lt;br /&gt;
wait_time_ref = {}&lt;br /&gt;
node = 0&lt;br /&gt;
while len(individuals) != 1:&lt;br /&gt;
    node_dict[node] = {&amp;quot;child_nodes&amp;quot;: []}&lt;br /&gt;
    sample_events = tuple(sample(individuals, 2))&lt;br /&gt;
    reference[sample_events] = node&lt;br /&gt;
    wait_time = exp(1/comb(len(individuals), 2))&lt;br /&gt;
    wait_time_ref[node] = wait_time&lt;br /&gt;
    for sample_event in sample_events:&lt;br /&gt;
        if type(sample_event) == str:&lt;br /&gt;
            node_dict[node][&amp;quot;child_nodes&amp;quot;].append(sample_event)&lt;br /&gt;
        else:&lt;br /&gt;
            child_node = reference[sample_event]&lt;br /&gt;
            node_dict[node][&amp;quot;child_nodes&amp;quot;].append(child_node)&lt;br /&gt;
    individuals.remove(sample_events[0])&lt;br /&gt;
    individuals.remove(sample_events[1])&lt;br /&gt;
    individuals.append(sample_events)&lt;br /&gt;
    node += 1&lt;br /&gt;
    &lt;br /&gt;
# Calculate Tree Branch Lengths&lt;br /&gt;
tree_dict = {}&lt;br /&gt;
cumulative_br_length = 0&lt;br /&gt;
cumulative_br_len_dict = {}&lt;br /&gt;
for i in range(len(node_dict)):&lt;br /&gt;
    tree_dict[i] = {}&lt;br /&gt;
    cumulative_br_length += wait_time_ref[i]&lt;br /&gt;
    cumulative_br_len_dict[i] = cumulative_br_length&lt;br /&gt;
    for node in node_dict[i][&#039;child_nodes&#039;]:&lt;br /&gt;
        if type(node) == str:&lt;br /&gt;
            tree_dict[i][node] = cumulative_br_length&lt;br /&gt;
        else:&lt;br /&gt;
            tree_dict[i][node] = cumulative_br_len_dict[i] - cumulative_br_len_dict[node]&lt;br /&gt;
            &lt;br /&gt;
# Parse the Tree into a String&lt;br /&gt;
for i in range(len(tree_dict)):&lt;br /&gt;
    for node in tree_dict[i].keys():&lt;br /&gt;
        if type(node) != str:&lt;br /&gt;
            temp = str(tree_dict[node])&lt;br /&gt;
            tree_dict[i][temp] = tree_dict[i].pop(node)&lt;br /&gt;
tree_str = str(tree_dict[i]).replace(&#039;\&#039;&#039;, &#039;&#039;).replace(&#039;\&amp;quot;&#039;, &#039;&#039;).replace(&#039;\\&#039;, &#039;&#039;).replace(&#039;{&#039;, &#039;(&#039;).replace(&#039;}&#039;, &#039;)&#039;).replace(&#039; &#039;, &#039;&#039;)&lt;br /&gt;
tree_str += &amp;quot;;&amp;quot;            &lt;br /&gt;
&lt;br /&gt;
# Visualize Tree&lt;br /&gt;
handle = StringIO(tree_str)&lt;br /&gt;
tree = Phylo.read(handle, &#039;newick&#039;)&lt;br /&gt;
tree.ladderize()   # Flip branches so deeper clades are displayed at top&lt;br /&gt;
Phylo.draw(tree)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==April 17, 2017 &amp;quot;Mutation Meltdown&amp;quot; (Due May 5, 2017)==&lt;br /&gt;
[[File:Ratchet-1.png|thumbnail]]&lt;br /&gt;
Our last exploration showed that population will reach a steady-state level of DNA sequence polymorphism under the opposing forces of genetic drift and mutations. The steady-state level is expected to be ~ N * mu: the larger the population size and the higher the mutation rate, the higher per-site DNA polymorphism. (Also note that, given a long enough sequence and a low enough mutation rate, we don&#039;t expect to see two or more mutations hitting a single position. Each mutation is essentially a new one and only two-state SNPs are expected in a sample of DNA sequences).&lt;br /&gt;
&lt;br /&gt;
However, the steady-state expectation is based on the assumption that all mutations are neutral (i.e., no beneficial or harmful fitness effects). In reality, mutations are predominantly either neutral or harmful and few are beneficial. Natural selection (negative or positive, except those on the immune-defense loci) tends to drive down genetic variation.&lt;br /&gt;
&lt;br /&gt;
Sex to the rescue. Without sex or recombination, the fate of genetic variations across a genome are bundled together, rising or sinking in unison with the fate of a single beneficial or harmful mutation. With recombination, genetic variations at different loci become less tightly linked and are more likely maintained, speeding up adaptation.&lt;br /&gt;
&lt;br /&gt;
This week, we will use simulation to recreate the so-called &amp;quot;[https://en.wikipedia.org/wiki/Muller%27s_ratchet Muller&#039;s Ratchet]&amp;quot;, which predicts that asexual populations are evolutionary dead ends. It shows that the combined effect of deleterious mutation and genetic drift will lead to population extinction due to steady accumulation of deleterious mutations, a phenomenon called &amp;quot;mutation meltdown&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
The simulation protocol would be similar to that of the last problem involving only mutation and drift, but I suggest you rewrite from scratch to be more efficient by using a simpler data structure (no sequence or bases needed). The main difference is that, instead of calculating average pairwise sequence differences in each generation, you will track the frequency of mutation-free sequences (n0) and find the time until it goes to zero. That is when the Ratchet makes a &amp;quot;click&amp;quot;. The next click is when the population losses the one-mutation sequences (n1), and so on. Each click drives the population fitness down by one mutation-level. [ Numerically, if each mutation causes a fitness loss of &amp;lt;code&amp;gt;s&amp;lt;/code&amp;gt; (&amp;quot;selection coefficient&amp;quot;), the fitness of a sequence with &amp;lt;code&amp;gt;k&amp;lt;/code&amp;gt; mutations is given by w = (1-s)&amp;lt;sup&amp;gt;k&amp;lt;/sup&amp;gt;. Strictly speaking, the selection coefficient should be included to affect gamete size, but let&#039;s ignore that for now].&lt;br /&gt;
&lt;br /&gt;
If there is recombination, the mutation-free sequences could be recovered (simulation next time?). Without recombination, the only direction for the population to evolve is a steady loss of best-fit individuals (by genetic drift).&lt;br /&gt;
&lt;br /&gt;
Questions:&lt;br /&gt;
# Does the ratchet occur faster in small or large populations? Find by simulating N=100 and N=1000&lt;br /&gt;
# Does the ratchet occur faster for a short or long genome? Find by simulating L=1e3 and L=1e4&lt;br /&gt;
# Which parts of the human genomes are asexual (therefore subject to Muller&#039;s Ratchet and becoming increasingly small)?&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
|&lt;br /&gt;
*By John&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;python&amp;quot;&amp;gt;&lt;br /&gt;
import numpy as np&lt;br /&gt;
import matplotlib.pyplot as plt&lt;br /&gt;
from numpy.random import poisson&lt;br /&gt;
from numpy.random import choice&lt;br /&gt;
&lt;br /&gt;
def simulator(seq_length, pop, repro_rate, generations):&lt;br /&gt;
    mutation_rate = 0.00001&lt;br /&gt;
    lam = seq_length * mutation_rate&lt;br /&gt;
    individuals = np.array([1 for i in range(pop)])&lt;br /&gt;
    non_mutated_pop_rate = []&lt;br /&gt;
    for generation in range(generations):&lt;br /&gt;
        gametes = []&lt;br /&gt;
        for individual in individuals:&lt;br /&gt;
            off_sprs = poisson(lam, repro_rate)&lt;br /&gt;
            mutation_ix = np.array(np.where(off_sprs != 0))&lt;br /&gt;
            indiv_gametes = np.array([individual for j in range(repro_rate)])&lt;br /&gt;
            indiv_gametes[mutation_ix] = 0&lt;br /&gt;
            gametes += list(indiv_gametes)&lt;br /&gt;
        individuals = choice(gametes, size=pop)&lt;br /&gt;
        non_mutated_frequency = float(np.count_nonzero(individuals)) / pop&lt;br /&gt;
        non_mutated_pop_rate.append(non_mutated_frequency)&lt;br /&gt;
    return non_mutated_pop_rate&lt;br /&gt;
&lt;br /&gt;
genome_length_1000 = 1000&lt;br /&gt;
genome_length_10000 = 10000&lt;br /&gt;
results_1000 = simulator(genome_length_1000, 1000, 100, 500)&lt;br /&gt;
results_10000 = simulator(genome_length_10000, 1000, 100, 500)&lt;br /&gt;
&lt;br /&gt;
genome_length_1000 = 1000&lt;br /&gt;
genome_length_10000 = 10000&lt;br /&gt;
results2_1000 = simulator(genome_length_1000, 10000, 100, 500)&lt;br /&gt;
results2_10000 = simulator(genome_length_10000, 10000, 100, 500)&lt;br /&gt;
&lt;br /&gt;
f, (ax1, ax2) = plt.subplots(1, 2, sharey=True)&lt;br /&gt;
line_1_1000 = ax1.plot(results_1000, &amp;quot;r&amp;quot;)&lt;br /&gt;
line_1_10000 = ax1.plot(results_10000, &amp;quot;b&amp;quot;)&lt;br /&gt;
line_2_1000 = ax2.plot(results2_1000, &amp;quot;r&amp;quot;)&lt;br /&gt;
line_2_10000 = ax2.plot(results2_10000, &amp;quot;b&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
ax1.set_title(&amp;quot;Population = 1000&amp;quot;)&lt;br /&gt;
ax2.set_title(&amp;quot;Population = 10000&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
plt.title(&amp;quot;Mutation Meltdown&amp;quot;)&lt;br /&gt;
plt.show()&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Weigang&lt;br /&gt;
&amp;lt;syntaxhighlight lang=R&amp;quot;&amp;gt;&lt;br /&gt;
add.mutation &amp;lt;- function(genome, genome.length, mutation.rate) {&lt;br /&gt;
  mu.exp &amp;lt;- genome.length * mutation.rate;&lt;br /&gt;
  mu.num &amp;lt;- rpois(1, lambda = mu.exp);&lt;br /&gt;
  return(genome+mu.num);&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
ratchet &amp;lt;- function(pop.size=100, gamete.size=100, mu=1e-4, genome.length=1e4) {&lt;br /&gt;
  out &amp;lt;- data.frame(generation=numeric(), n0=numeric(), n1=numeric(), pop=numeric(), length=numeric());&lt;br /&gt;
  pop &amp;lt;- rep(0, pop.size) # initial all mutation-free&lt;br /&gt;
  g &amp;lt;- 1; # generation&lt;br /&gt;
  freq0 &amp;lt;- 1; # freq of zero-class&lt;br /&gt;
  while(freq0 &amp;gt; 0) {&lt;br /&gt;
    cat(&amp;quot;at generation&amp;quot;, g, &amp;quot;\n&amp;quot;);&lt;br /&gt;
    freq0 &amp;lt;- length(which(pop == 0));&lt;br /&gt;
    out &amp;lt;- rbind(out, data.frame(generation=g, n0=freq0/pop.size, pop=pop.size, length=genome.length));&lt;br /&gt;
    gametes &amp;lt;- sapply(1:gamete.size, function(x) {lapply(pop, function(x) add.mutation(x, genome.length, mu))});&lt;br /&gt;
    pop &amp;lt;- sample(gametes, pop.size);&lt;br /&gt;
    g &amp;lt;- g+1;&lt;br /&gt;
  }&lt;br /&gt;
  return(out);&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
out.long.genome.df &amp;lt;- ratchet(genome.length=1e4);&lt;br /&gt;
out.short.genome.df &amp;lt;- ratchet(genome.length=1e3);&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==March 31, 2017 &amp;quot;Drift-Mutation Balance&amp;quot; (4/14/2017)==&lt;br /&gt;
[[File:Drift-mutation.png|thumbnail]]&lt;br /&gt;
We will explore genetic drift of a DNA fragment with mutation under Wright-Fisher model. From last week&#039;s exercise, we conclude that a population will sooner or later lose genetic diversity (becoming fixed after ~2N generations), if no new alleles are generated (by e.g., mutation or migration). &lt;br /&gt;
&lt;br /&gt;
Mutation, in contrast, increases genetic diversity over time. Under neutrality (no natural selection against or for any mutation, e.g., on an intron sequence), the population will reach an equilibrium point when the loss of genetic diversity by drift is cancelled out by increase of genetic diversity by mutation.&lt;br /&gt;
&lt;br /&gt;
You job is to find this equilibrium point by simulation, given a population size (N) and a mutation rate (mu). The expected answer is pi=2N*mu, where pi is a measure of genetic diversity using DNA sequences, which is the average pairwise sequence differences within a population.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;font color=&amp;quot;blue&amp;quot;&amp;gt;Note: the following algorithm seems to be too complex to build at once. Also, too slow to run in R. Nonetheless, please try to write two R functions: (1)mutate.seq(seq, mut), with &amp;quot;seq&amp;quot; as a vector of bases and &amp;quot;mut&amp;quot; is the mutation rate, and (2) avg.seq.diff(pop), with &amp;quot;pop&amp;quot; as a list of DNA sequences&amp;lt;/font&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A suggested algorithm is:&lt;br /&gt;
# Start with a homogeneous population with N=100 identical DNA sequences (e.g., with a length of L=1e4 bases, or about 5 genes) [R hint: &amp;lt;code&amp;gt;dna &amp;lt;- sample(c(&amp;quot;a&amp;quot;, &amp;quot;t&amp;quot;, &amp;quot;c&amp;quot;, &amp;quot;g&amp;quot;), size=1e5, replace=T, prob = rep(0.25,4))&amp;lt;/code&amp;gt;]&lt;br /&gt;
# Write a mutation function, which will mutate the DNA based on Poisson process (since mutation is a rare event). For example, if mu=1e-4 per generation per base per individual (too high for real, but faster for results to converge), then each generation the expected number of mutations would be L * mu = 1 per individual per generation for our DNA segment. You would then simulate the random number of mutations by using the R function &amp;lt;code&amp;gt;num.mutations &amp;lt;- rpois(1,lambda=1)&amp;lt;/code&amp;gt;.&lt;br /&gt;
# Apply the mutation function for each individual DNA copy (a total of N=100) during gamete production (100 gametes for each individual) at each generation (for a total of G=1000 generations).&lt;br /&gt;
# Write another function to calculate, for each generation, instead of counting allele frequencies (as last week&#039;s problem), to calculate &amp;amp; output average pairwise differences among the 100 individuals.&lt;br /&gt;
# Finally, you would graph pi over generation.&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
| &lt;br /&gt;
* By John&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;python&amp;quot;&amp;gt;&lt;br /&gt;
import numpy as np&lt;br /&gt;
import matplotlib.pyplot as plt&lt;br /&gt;
from itertools import combinations&lt;br /&gt;
from random import sample&lt;br /&gt;
from numpy.random import choice&lt;br /&gt;
from numpy.random import poisson&lt;br /&gt;
&lt;br /&gt;
# global variables&lt;br /&gt;
nuc = np.array([&amp;quot;A&amp;quot;, &amp;quot;T&amp;quot;, &amp;quot;C&amp;quot;, &amp;quot;G&amp;quot;])&lt;br /&gt;
expected_mutation_rate = 0.00001 # there are 2 sequences per individual&lt;br /&gt;
seq_length = 10000&lt;br /&gt;
individual_pop = 100&lt;br /&gt;
gamete_rate = 100&lt;br /&gt;
generation = 1000&lt;br /&gt;
&lt;br /&gt;
def simulator(population, repro_rate, generation):&lt;br /&gt;
    observed_mutation_collection = []&lt;br /&gt;
    &lt;br /&gt;
    # Original Individuals&lt;br /&gt;
    original = choice(nuc, seq_length, 0.25)&lt;br /&gt;
    individual_total = [original for i in range(population)]&lt;br /&gt;
    &lt;br /&gt;
    for i in range(generation): # iterate over 1000 generations&lt;br /&gt;
        # Produce Gametes&lt;br /&gt;
        gametes = []&lt;br /&gt;
        for individual in individual_total:&lt;br /&gt;
            gametes += [individual for i in range(repro_rate)]&lt;br /&gt;
        gametes = np.array(gametes)&lt;br /&gt;
        gametes_pop = gametes.shape[0] # number of gametes: 100 * 100&lt;br /&gt;
        &lt;br /&gt;
        # Mutation&lt;br /&gt;
        mutation_number_arr = poisson(lam=seq_length * expected_mutation_rate, size=gametes_pop) # derive number of mutations per individual&lt;br /&gt;
        mutation_index_arr = [sample(range(seq_length), mutations) for mutations in mutation_number_arr] # get the index of mutation&lt;br /&gt;
        # Mutation: Replace with Mututated Base Pairs&lt;br /&gt;
        for i in range(gametes_pop): # iterate over all gametes&lt;br /&gt;
            for ix in mutation_index_arr[i]: # iterate over mutated base pair for each gamete&lt;br /&gt;
                gametes[i, ix] = choice(nuc[nuc != gametes[i, ix]], 1)[0] # locate mutation and alter the nuc with mutated one&lt;br /&gt;
        &lt;br /&gt;
        # Next generation of individuals&lt;br /&gt;
        individual_total = gametes[choice(range(10000), 100, replace=False)]&lt;br /&gt;
        &lt;br /&gt;
        # Calculate observed mutation rate&lt;br /&gt;
        num_combinations = 0&lt;br /&gt;
        total_diff_bases = 0&lt;br /&gt;
        for pair in combinations(range(len(individual_total)), 2): # aggregate mutations for all possible pairs of individuals&lt;br /&gt;
            total_diff_bases += len(np.where((individual_total[pair[0]] == individual_total[pair[1]]) == False)[0])&lt;br /&gt;
            num_combinations += 1&lt;br /&gt;
        observed_mutation_rate = float(total_diff_bases) / float(num_combinations) / seq_length&lt;br /&gt;
        observed_mutation_collection.append(observed_mutation_rate)&lt;br /&gt;
        &lt;br /&gt;
    return(observed_mutation_collection)&lt;br /&gt;
&lt;br /&gt;
# Simulation&lt;br /&gt;
result = simulator(100, 100, 1000)&lt;br /&gt;
&lt;br /&gt;
# Visualize results&lt;br /&gt;
plt.plot(result, &#039;b&#039;)&lt;br /&gt;
plt.xlabel(&amp;quot;Generation&amp;quot;)&lt;br /&gt;
plt.ylabel(&amp;quot;Mutation Rate&amp;quot;)&lt;br /&gt;
plt.title(&amp;quot;Drift-Mutation Balance&amp;quot;)&lt;br /&gt;
plt.show()&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt; &lt;br /&gt;
&lt;br /&gt;
*By Weigang&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
library(Biostrings) # a more compressed way to store and manipulate sequences&lt;br /&gt;
bases &amp;lt;- DNA_ALPHABET[1:4];&lt;br /&gt;
dna &amp;lt;- sample(bases, size = 1e5, replace = T);&lt;br /&gt;
library(ape) # to use the DNAbin methods&lt;br /&gt;
# function to mutate (Poisson process)&lt;br /&gt;
mutate.seq &amp;lt;- function(seq, mutation.rate) {&lt;br /&gt;
  mu.exp &amp;lt;- length(seq) * mutation.rate;&lt;br /&gt;
  mu.num &amp;lt;- rpois(1, lambda = mu.exp);&lt;br /&gt;
  if (mu.num &amp;gt; 0) {&lt;br /&gt;
    pos &amp;lt;- sample(1:length(seq), size = mu.num);&lt;br /&gt;
    for (j in 1:length(pos)) {&lt;br /&gt;
      current.base &amp;lt;- seq[pos[j]];&lt;br /&gt;
      mutated.base &amp;lt;-sample(bases[bases !=current.base)], size = 1);&lt;br /&gt;
      seq[pos[j]] &amp;lt;- mutated.base;&lt;br /&gt;
    }&lt;br /&gt;
  }&lt;br /&gt;
  return(seq)&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
# main routine&lt;br /&gt;
drift.mutation &amp;lt;- function(pop.size=100, generation.time=500, gamete.size=10, mu=1e-5) {&lt;br /&gt;
  out &amp;lt;- data.frame(generation=numeric(), pi=numeric()); # for storing outputs&lt;br /&gt;
  pop &amp;lt;- lapply(1:pop.size, function(x) dna)  # create the initial population&lt;br /&gt;
  for(i in 1:generation.time) { # for each generation&lt;br /&gt;
    cat(&amp;quot;at generation&amp;quot;, i, &amp;quot;\n&amp;quot;); # print progress&lt;br /&gt;
    gametes &amp;lt;- sapply(1:gamete.size, function(x) {lapply(pop, function(x) mutate.seq(x, mu))}); &lt;br /&gt;
    pop &amp;lt;- sample(gametes, pop.size); &lt;br /&gt;
    pop.bin &amp;lt;- as.DNAbin(pop); # change into DNAbin class to advantage of its dist.dna() function&lt;br /&gt;
    out &amp;lt;- rbind(out, data.frame(generation=i, pi=mean(dist.dna(pop.bin))));&lt;br /&gt;
  }&lt;br /&gt;
  out;&lt;br /&gt;
}&lt;br /&gt;
out.df &amp;lt;- drift.mutation(); # run function and save results into a data frame&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==March 18, 2017 &amp;quot;Genetic Drift&amp;quot; (Due 3/31/2017)==&lt;br /&gt;
[[File:Drift.png|thumbnail]]&lt;br /&gt;
[[File:Drift-founder.png|thumbnail]]&lt;br /&gt;
This is our first biological simulation. Mandatory assignment for all lab members (from interns to doctoral students). An expected result is shown in the graph.&lt;br /&gt;
&lt;br /&gt;
Task: Simulate the Wright-Fisher model of genetic drift as follows:&lt;br /&gt;
# Begin with an allele frequency p=0.5, and a pop of N=100 haploid individuals [Hint: &amp;lt;code&amp;gt;pop &amp;lt;- c(rep(&amp;quot;A&amp;quot;, 50), rep(&amp;quot;a&amp;quot;, 50))&amp;lt;/code&amp;gt;]&lt;br /&gt;
# Each individual produces 100 gametes, giving a total of 10,000 gametes [Hint: use a for loop with rep() function]&lt;br /&gt;
# Sample from the gamete pool another 100 to give rise to a new generation of individuals [Hint: use sample() function]&lt;br /&gt;
# Calculate allele frequency [Hint: use table() function]&lt;br /&gt;
# Repeat the above in succession for a total generation of g=1000 generations [Hint: create a function with three arguments, e.g., wright.fisher(pop.size, gamete.size, generation.time)]&lt;br /&gt;
# Plot allele frequency changes over generation time&lt;br /&gt;
# Be prepared to answer these questions:&lt;br /&gt;
## Why allele frequency fluctuate even without natural selection?&lt;br /&gt;
## What&#039;s the final fate of population, one allele left, or two alleles coexist indefinitely?&lt;br /&gt;
## Which population can maintain genetic polymorphism (with two alleles) longer?&lt;br /&gt;
## Which population gets fixed (only one allele remains) quicker?&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
|&lt;br /&gt;
* By Lili (May, 2019)&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
#Simple sampling, haploid, as a function &lt;br /&gt;
wright_fisher &amp;lt;- function(pop_size, gam_size, alle_frq, n_gen) {&lt;br /&gt;
  pop &amp;lt;- c(rep(&amp;quot;A&amp;quot;, pop_size*frq), rep(&amp;quot;a&amp;quot;, pop_size*frq))&lt;br /&gt;
  prob &amp;lt;- numeric(n_gen)&lt;br /&gt;
  for(time in 1:n_gen){&lt;br /&gt;
    gamt &amp;lt;- rep(pop, gam_size)&lt;br /&gt;
    pop &amp;lt;- sample(gamt, pop_size)&lt;br /&gt;
    prob[time] &amp;lt;- table(pop)[1]/pop_size&lt;br /&gt;
  }&lt;br /&gt;
  wf_df &amp;lt;- data.frame(generation=1:n_gen, probability=prob)&lt;br /&gt;
  plot(drift_df, type= &#039;l&#039;, main = &amp;quot;Wright-Fisher Process (Genetic Drift)&amp;quot;, xlab = &amp;quot;generation&amp;quot;, ylab = &amp;quot;allele frequency&amp;quot;)&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
wright_fisher(1000, 100, 0.5, 1000)&lt;br /&gt;
&lt;br /&gt;
#Version 2. Using binomial distribution with replicates, diploid&lt;br /&gt;
&lt;br /&gt;
pop_size &amp;lt;- c(50, 100, 1000, 5000)&lt;br /&gt;
alle_frq &amp;lt;- c(0.01, 0.1, 0.5, 0.8)&lt;br /&gt;
n_gen &amp;lt;- 100&lt;br /&gt;
n_reps &amp;lt;- 50&lt;br /&gt;
genetic_drift &amp;lt;- data.frame()&lt;br /&gt;
&lt;br /&gt;
for(N in pop_size){&lt;br /&gt;
  for(p in alle_frq){&lt;br /&gt;
    p0 &amp;lt;- p&lt;br /&gt;
    for(j in 1:n_gen){&lt;br /&gt;
      X &amp;lt;- rbinom(n_reps, 2*N, p)&lt;br /&gt;
      p &amp;lt;- X/(2*N)&lt;br /&gt;
      rows &amp;lt;- data.frame(replicate= 1:n_reps, pop=rep(N, n_reps), gen=rep(j, n_reps), frq=rep(p0, n_reps), prob=p )&lt;br /&gt;
      genetic_drift &amp;lt;- rbind(genetic_drift, rows)&lt;br /&gt;
    }&lt;br /&gt;
  }&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
library(ggplot2)&lt;br /&gt;
ggplot(genetic_drift, aes(x=gen, y=prob, group=replicate)) + geom_path(alpha= .5) + facet_grid(pop ~ frq) + guides(colour=FALSE)&lt;br /&gt;
&lt;br /&gt;
# 3rd version: trace ancestry (instead of allele frequency)&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By John&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;python&amp;quot;&amp;gt;&lt;br /&gt;
#Python&lt;br /&gt;
import numpy as np&lt;br /&gt;
import sys&lt;br /&gt;
from random import sample&lt;br /&gt;
import matplotlib.pyplot as plt&lt;br /&gt;
 &lt;br /&gt;
def simulator(gametes_rate, next_individuals, generations):&lt;br /&gt;
    individuals = [1 for i in range(50)] + [0 for j in range(50)]&lt;br /&gt;
    frequency = []&lt;br /&gt;
    for generation in range(generations):&lt;br /&gt;
        gametes = []&lt;br /&gt;
        for individual in individuals:&lt;br /&gt;
            gametes += [individual for i in range(gametes_rate)]&lt;br /&gt;
        individuals = sample(gametes, next_individuals)&lt;br /&gt;
        frequency.append(np.count_nonzero(individuals) / len(individuals))&lt;br /&gt;
    return(frequency)&lt;br /&gt;
 &lt;br /&gt;
N_100 = simulator(100, 100, 1000)&lt;br /&gt;
N_1000 = simulator(100, 1000, 1000)&lt;br /&gt;
 &lt;br /&gt;
# Create plots with pre-defined labels.&lt;br /&gt;
# Alternatively,pass labels explicitly when calling `legend`.&lt;br /&gt;
fig, ax = plt.subplots()&lt;br /&gt;
ax.plot(N_100, &#039;r&#039;, label=&#039;N=100&#039;)&lt;br /&gt;
ax.plot(N_1000, &#039;b&#039;, label=&#039;N=1000&#039;)&lt;br /&gt;
 &lt;br /&gt;
# Add x, y labels and title&lt;br /&gt;
plt.ylim(-0.1, 1.1)&lt;br /&gt;
plt.xlabel(&amp;quot;Generation&amp;quot;)&lt;br /&gt;
plt.ylabel(&amp;quot;Frequency&amp;quot;)&lt;br /&gt;
plt.title(&amp;quot;Wright-Fisher Model&amp;quot;)&lt;br /&gt;
 &lt;br /&gt;
# Now add the legend with some customizations.&lt;br /&gt;
legend = ax.legend(loc=&#039;upper right&#039;, shadow=True)&lt;br /&gt;
 &lt;br /&gt;
# The frame is matplotlib.patches.Rectangle instance surrounding the legend.&lt;br /&gt;
frame = legend.get_frame()&lt;br /&gt;
frame.set_facecolor(&#039;0.90&#039;)&lt;br /&gt;
 &lt;br /&gt;
# Set the fontsize&lt;br /&gt;
for label in legend.get_texts():&lt;br /&gt;
    label.set_fontsize(&#039;large&#039;)&lt;br /&gt;
 &lt;br /&gt;
for label in legend.get_lines():&lt;br /&gt;
    label.set_linewidth(1.5)  # the legend line width&lt;br /&gt;
plt.show()&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;scala&amp;quot;&amp;gt;&lt;br /&gt;
#Scala&lt;br /&gt;
import scala.util.Random&lt;br /&gt;
import scala.collection.mutable.ListBuffer&lt;br /&gt;
 &lt;br /&gt;
val gamete_rate = 100&lt;br /&gt;
val offspr_rate = 100&lt;br /&gt;
val generations = 1000&lt;br /&gt;
 &lt;br /&gt;
var frequency: List[Double] = List()&lt;br /&gt;
 &lt;br /&gt;
var individuals = List.fill(50)(0) ++ List.fill(50)(1)&lt;br /&gt;
 &lt;br /&gt;
for(generation &amp;lt;- 1 to generations){&lt;br /&gt;
  var gametes = individuals.map(x =&amp;gt; List.fill(gamete_rate)(x)).flatten&lt;br /&gt;
  individuals = Random.shuffle(gametes).take(offspr_rate)&lt;br /&gt;
  frequency = frequency :+ individuals.count(_ == 1).toDouble / offspr_rate&lt;br /&gt;
}&lt;br /&gt;
 &lt;br /&gt;
print(frequency)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;java&amp;quot;&amp;gt;&lt;br /&gt;
#Spark&lt;br /&gt;
val gamete_rate = 100&lt;br /&gt;
val offspr_rate = 100&lt;br /&gt;
val generations = 300&lt;br /&gt;
 &lt;br /&gt;
var frequency: List[Double] = List()&lt;br /&gt;
var individuals = sc.parallelize(List.fill(50)(&amp;quot;0&amp;quot;) ++ List.fill(50)(&amp;quot;1&amp;quot;))&lt;br /&gt;
 &lt;br /&gt;
for(generation &amp;lt;- 1 to generations){&lt;br /&gt;
  val gametes = individuals.flatMap(x =&amp;gt; (x * gamete_rate).split(&amp;quot;&amp;quot;).tail)&lt;br /&gt;
  individuals = sc.parallelize(gametes.takeSample(false, offspr_rate))&lt;br /&gt;
  val count = individuals.countByValue&lt;br /&gt;
  frequency = frequency :+ count(&amp;quot;1&amp;quot;).toDouble / offspr_rate&lt;br /&gt;
}&lt;br /&gt;
 &lt;br /&gt;
print(frequency)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Brian&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
Wright&amp;lt;-function(pop.size,gam.size,generation) {&lt;br /&gt;
 if( pop.size%%2==0) {&lt;br /&gt;
 &lt;br /&gt;
  pop&amp;lt;-c(rep(&amp;quot;A&amp;quot;,pop.size*.5),rep(&amp;quot;a&amp;quot;,pop.size))&lt;br /&gt;
  frequency&amp;lt;-.5&lt;br /&gt;
  time&amp;lt;-0&lt;br /&gt;
  geneticDrift&amp;lt;-data.frame(frequency,time)&lt;br /&gt;
  for (i in 1:generation) {&lt;br /&gt;
    largePop&amp;lt;-rep(pop,gam.size)&lt;br /&gt;
    samplePop&amp;lt;-sample(largePop,100)&lt;br /&gt;
    cases&amp;lt;-table(samplePop)&lt;br /&gt;
    if (cases[1]&amp;lt;100 &amp;amp;&amp;amp; cases[2]&amp;lt;100 ) &lt;br /&gt;
        {propA&amp;lt;-(cases[1]/100)&lt;br /&gt;
        geneticDrift&amp;lt;-rbind(geneticDrift,c(propA,i))&lt;br /&gt;
        pop&amp;lt;-c(rep(&amp;quot;A&amp;quot;,cases[2]),rep(&amp;quot;a&amp;quot;,cases[1]))&lt;br /&gt;
  }  &lt;br /&gt;
}&lt;br /&gt;
  plot(geneticDrift$frequency~geneticDrift$time,type=&amp;quot;b&amp;quot;, main=&amp;quot;Genetic Drift N=1000&amp;quot;, xlab=&amp;quot;time&amp;quot;,ylab=&amp;quot;Proportion Pop A&amp;quot;,col=&amp;quot;red&amp;quot;, pch=18 )&lt;br /&gt;
}&lt;br /&gt;
 if(pop.size%%2==1 ) {&lt;br /&gt;
   print(&amp;quot;Initial Population should be even. Try again&amp;quot;)&lt;br /&gt;
 }&lt;br /&gt;
}&lt;br /&gt;
Wright(1000,100,1000)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Jamila&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
Genetic_code = function(t,R){ &lt;br /&gt;
  N&amp;lt;- 100&lt;br /&gt;
  p&amp;lt;- 0.5&lt;br /&gt;
  frequency&amp;lt;-as.numeric();&lt;br /&gt;
  for (i in 1:t){&lt;br /&gt;
    A1=rbiom(1,2*N,p)&lt;br /&gt;
    p=A1/(N*2); &lt;br /&gt;
    frequency[length(frequency)+1]&amp;lt;-p; &lt;br /&gt;
  }&lt;br /&gt;
  plot(frequency, type=&amp;quot;1&amp;quot;,ylim=c(0,1),col=3,xlab=&amp;quot;t&amp;quot;,ylab=expression(p(A[1])))&lt;br /&gt;
  &lt;br /&gt;
  }&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Sharon&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
p=0.5&lt;br /&gt;
N=100&lt;br /&gt;
g=0&lt;br /&gt;
pop &amp;lt;- c(rep(&amp;quot;A&amp;quot;, 50), rep(&amp;quot;a&amp;quot;, 50)) #create a population of 2 alleles&lt;br /&gt;
data&amp;lt;-data.frame(g,p) #a set of variables of the same # of rows &lt;br /&gt;
for (i in 1:1000) { #generation &lt;br /&gt;
  gam_pl &amp;lt;-rep(pop, 100) #gamete pool&lt;br /&gt;
  gam_sam &amp;lt;-sample(gam_pl, 100)  #sample from gamete pool&lt;br /&gt;
  tab_all&amp;lt;-table(gam_sam) #table ps sample&lt;br /&gt;
  all_freq&amp;lt;-tab_all[1]/100 #get allele frequency&lt;br /&gt;
  data&amp;lt;-rbind(data, c(i,all_freq))&lt;br /&gt;
  pop&amp;lt;-c(rep(&amp;quot;A&amp;quot;,tab_all[1]))&lt;br /&gt;
}&lt;br /&gt;
plot(data$g,data$p)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Sipa&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
pop &amp;lt;- c(rep(&amp;quot;A&amp;quot;, 50), rep(&amp;quot;a&amp;quot;, 50))&lt;br /&gt;
wright_fisher &amp;lt;- function(N,gen_time,gam) {&lt;br /&gt;
  N=N&lt;br /&gt;
  gen_time = gen_time&lt;br /&gt;
  x = numeric(gen_time) &lt;br /&gt;
  x[1] = gam&lt;br /&gt;
  for (i in 2:1000) {&lt;br /&gt;
  k=(x[i-1])/N&lt;br /&gt;
  n=seq(0,N,1)&lt;br /&gt;
  prob=dbinom(n,N,k)&lt;br /&gt;
  x[i]=sample(0:N, 1, prob=prob)&lt;br /&gt;
  &lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
plot(x[1:gen_time], type=&amp;quot;l&amp;quot;, pch=10)&lt;br /&gt;
&lt;br /&gt;
}&lt;br /&gt;
pool &amp;lt;- rep(pop, times = 100)&lt;br /&gt;
s_pool &amp;lt;- sample(pool, size = 100, replace = F)&lt;br /&gt;
table(s_pool)&lt;br /&gt;
wright_fisher2 &amp;lt;- function(pop_size,al_freq,gen_time) {&lt;br /&gt;
  pop &amp;lt;- c(rep(&amp;quot;A&amp;quot;, pop_size/2), rep(&amp;quot;a&amp;quot;, pop_size/2))&lt;br /&gt;
  pool &amp;lt;-rep(pop, times=100)&lt;br /&gt;
  s_pool &amp;lt;- sample(pool, size = 100, replace = T)&lt;br /&gt;
  for(i in 1:gen_time)&lt;br /&gt;
    s_pool &amp;lt;- sample(sample(pool, size = pop_size, replace = F))&lt;br /&gt;
  a.f &amp;lt;- table(s_pool)[1]/100&lt;br /&gt;
  return(a.f)&lt;br /&gt;
  }&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Nicolette&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
N=100 &lt;br /&gt;
p=0.5 &lt;br /&gt;
g=1000 &lt;br /&gt;
sim=100 &lt;br /&gt;
Genetic_drift=array(0, dim=c(g,sim))&lt;br /&gt;
Genetic_drift[1,]=rep(N*p,sim)&lt;br /&gt;
for(i in 1:sim) {&lt;br /&gt;
  for(j in 2:g){&lt;br /&gt;
    X[j,i]=rbinom(1,N,prob=X[j-1,i]/N)&lt;br /&gt;
  }&lt;br /&gt;
}&lt;br /&gt;
Genetic_drift=data.frame(X/N)&lt;br /&gt;
matplot(1:1000, (X/N), type=&amp;quot;l&amp;quot;,ylab=&amp;quot;allele_frequency&amp;quot;,xlab=&amp;quot;generations&amp;quot;)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* By Weigang&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
# output frequency:&lt;br /&gt;
wright.fisher &amp;lt;- function(pop.size, generation.time=1000, gametes.per.ind=100) {&lt;br /&gt;
  out.df &amp;lt;- data.frame(gen=numeric(), freq=numeric(), pop.size=numeric());&lt;br /&gt;
  pop &amp;lt;- c(rep(&amp;quot;A&amp;quot;, pop.size/2), rep(&amp;quot;a&amp;quot;, pop.size/2)); # initial pop&lt;br /&gt;
  for(i in 1:generation.time) {&lt;br /&gt;
    gametes &amp;lt;- unlist(lapply(pop, function(x) rep(x, gametes.per.ind)));&lt;br /&gt;
    pop &amp;lt;- sample(gametes, pop.size); # sample gamete pool without replacement&lt;br /&gt;
    freq.a &amp;lt;- table(pop)[1]/pop.size; # frequency of &amp;quot;a&amp;quot;&lt;br /&gt;
    out.df &amp;lt;- rbind(out.df, c(gen=i, freq=freq.a, pop=pop.size, rep=rep));&lt;br /&gt;
  }&lt;br /&gt;
  out.list;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
out.1e2&amp;lt;-wright.fisher(100)&lt;br /&gt;
out.1e3&amp;lt;-wright.fisher(1000)&lt;br /&gt;
# Make Figure 1 above&lt;br /&gt;
plot(out.1e2[,1], out.1e2[,2], type = &amp;quot;l&amp;quot;, las=1, ylim=c(0,1), xlab=&amp;quot;generation&amp;quot;, ylab=&amp;quot;allele frequency&amp;quot;, main=&amp;quot;Wright-Fisher Process (Genetic Drift)&amp;quot;, col=3, lwd=2)&lt;br /&gt;
lines(out.1e3[,1], out.1e3[,2], type = &amp;quot;l&amp;quot;, col=2, lwd=2)&lt;br /&gt;
legend(400, 0.2, c(&amp;quot;N=100&amp;quot;, &amp;quot;N=1,000&amp;quot;), lty=1, col=3:2, cex=0.75)&lt;br /&gt;
abline(h=0.5, col=&amp;quot;gray&amp;quot;, lty=2)&lt;br /&gt;
&lt;br /&gt;
# Second function to track founders (not frequency)&lt;br /&gt;
wright.fisher.2 &amp;lt;- function(pop.size, generation.time=1000, gametes.per.ind=100) {&lt;br /&gt;
  pop &amp;lt;- 1:100 # label founders&lt;br /&gt;
  out.list &amp;lt;- list(pop);&lt;br /&gt;
  for(i in 1:generation.time) {&lt;br /&gt;
    gametes &amp;lt;- unlist(lapply(pop, function(x) rep(x, gametes.per.ind)));&lt;br /&gt;
    pop &amp;lt;- sample(gametes, pop.size); # sample gamete pool without replacement&lt;br /&gt;
    out.list[[i+1]] &amp;lt;- pop;&lt;br /&gt;
  }&lt;br /&gt;
  out.list;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
out.1ist&amp;lt;-wright.fisher.2(100)&lt;br /&gt;
# plot loss of genetic diversity by tracking surviving founder lineages&lt;br /&gt;
plot(x=1:100, y=rep(1, 100), ylim=c(0,100), xlim=c(1,600), type=&amp;quot;n&amp;quot;, las=1, main=&amp;quot;Drift: Suvival of founder lineages\n(N=100)&amp;quot;, cex.main=0.75, xlab=&amp;quot;generation&amp;quot;, ylab=&amp;quot;founders&amp;quot;)&lt;br /&gt;
for (i in 1:600) { points(rep(i,100), 1:100, col=colors()[out.list[[i]]], cex=0.5) }&lt;br /&gt;
for (i in 1:600) {points(rep(i,100), out.list[[i]], cex=0.2, col=&amp;quot;yellow&amp;quot;) }&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== March 11, 2017 &amp;quot;Stoplights&amp;quot; (Due 3/18/2017)==&lt;br /&gt;
[[File:Stoplight-1.png|framed|right]]&lt;br /&gt;
* Source: Paul Nahin (2008). &amp;quot;Digital Dice&amp;quot;, Problem 18.&lt;br /&gt;
* Challenge: How many red lights, on average, will you have to wait for on your journey from a city block m streets and n avenues away from Belfer [with coordinates (1,1)]? (assuming equal probability for red and green lights)&lt;br /&gt;
* Note that one has only wait for green light when walking along either the north side of 69 Street or east side of 1st Avenue. On all other intersections, one can walk non-stop without waiting for green light by crossing in the other direction if a red light is on.&lt;br /&gt;
* Formulate your solution with the following steps:&lt;br /&gt;
# Start from (m+1,n+1) corner and end at (1,1) corner&lt;br /&gt;
# The average number of red lights for m=n=0 is zero&lt;br /&gt;
# Find the average number of red lights for m=n=1 by simulating the walk 100 times&lt;br /&gt;
# Increment m &amp;amp; n by 1 (but keep m=n), until m=n=1000&lt;br /&gt;
# Plot average number of red lights by m (or n).&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
| &lt;br /&gt;
*By John&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;python&amp;quot;&amp;gt;&lt;br /&gt;
import numpy as np&lt;br /&gt;
import matplotlib.pyplot as plt&lt;br /&gt;
import pandas as pd&lt;br /&gt;
from pandas import DataFrame&lt;br /&gt;
from random import sample&lt;br /&gt;
from itertools import combinations, permutations&lt;br /&gt;
from numpy import count_nonzero, array&lt;br /&gt;
 &lt;br /&gt;
def red_light(av, st):&lt;br /&gt;
    traffic_light = [&amp;quot;red_st&amp;quot;, &amp;quot;red_av&amp;quot;]&lt;br /&gt;
    total_cross = av + st - 1&lt;br /&gt;
    while av != 0 and st != 0:&lt;br /&gt;
        if sample(traffic_light, 1)[0] == &amp;quot;red_st&amp;quot;:&lt;br /&gt;
            av -= 1&lt;br /&gt;
        else:&lt;br /&gt;
            st -= 1&lt;br /&gt;
    rest_cross = av if av != 0 else st&lt;br /&gt;
    return(count_nonzero([sample([0, 1], 1) for corss in range(rest_cross)]))&lt;br /&gt;
 &lt;br /&gt;
df = DataFrame(0, index=np.arange(10), columns=np.arange(10))&lt;br /&gt;
 &lt;br /&gt;
simulation = 1000&lt;br /&gt;
for av in df.index:&lt;br /&gt;
    for st in df.index:&lt;br /&gt;
        df.loc[av, st] = sum(array([[red_light(av, st)] for n in range(simulation)])) / simulation&lt;br /&gt;
        &lt;br /&gt;
plt.imshow(df, cmap=&#039;hot&#039;, interpolation=&#039;nearest&#039;)&lt;br /&gt;
plt.xticks(np.arange(1, 10))&lt;br /&gt;
plt.yticks(np.arange(1, 10))&lt;br /&gt;
plt.title(&amp;quot;Average Numbers Waiting for Red Light&amp;quot;)&lt;br /&gt;
plt.xlabel(&amp;quot;Number of Av&amp;quot;)&lt;br /&gt;
plt.ylabel(&amp;quot;Number of St&amp;quot;)&lt;br /&gt;
plt.show()&lt;br /&gt;
 &lt;br /&gt;
df&lt;br /&gt;
 &lt;br /&gt;
# Recursive Function for Each Trial&lt;br /&gt;
&lt;br /&gt;
from sys import stdout&lt;br /&gt;
av = 30&lt;br /&gt;
st = 30&lt;br /&gt;
n_space = [0]&lt;br /&gt;
traffic_light = [&amp;quot;red_st&amp;quot;, &amp;quot;red_av&amp;quot;]&lt;br /&gt;
def route(av, st):&lt;br /&gt;
    if not av == st == 0:&lt;br /&gt;
        if sample(traffic_light, 1)[0] == &amp;quot;red_st&amp;quot;:&lt;br /&gt;
            if av == 0:&lt;br /&gt;
                stdout.write(&amp;quot;w_&amp;quot;)&lt;br /&gt;
                n_space[0] += 1&lt;br /&gt;
                route(av, st - 1)&lt;br /&gt;
            else:&lt;br /&gt;
                stdout.write(&amp;quot;\n&amp;quot; + &amp;quot; &amp;quot; * n_space[0] + &amp;quot;|&amp;quot;)&lt;br /&gt;
                route(av - 1, st)&lt;br /&gt;
        else:&lt;br /&gt;
            if st == 0:&lt;br /&gt;
                stdout.write(&amp;quot;\n&amp;quot; + &amp;quot; &amp;quot; * n_space[0] + &amp;quot;w&amp;quot; + &amp;quot;\n&amp;quot; + &amp;quot; &amp;quot; * n_space[0] + &amp;quot;|&amp;quot;)&lt;br /&gt;
                route(av - 1, st)&lt;br /&gt;
            else:&lt;br /&gt;
                stdout.write(&amp;quot;_&amp;quot;)&lt;br /&gt;
                n_space[0] += 1&lt;br /&gt;
                route(av, st - 1)&lt;br /&gt;
route(av, st)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Weigang&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;perl&amp;quot;&amp;gt;&lt;br /&gt;
#!/usr/bin/env perl&lt;br /&gt;
use strict;&lt;br /&gt;
use warnings;&lt;br /&gt;
&lt;br /&gt;
# Use recursion&lt;br /&gt;
my ($distx, $disty) = @ARGV;&lt;br /&gt;
my $num_red = 0; # keep red-light counts&lt;br /&gt;
print &amp;quot;-&amp;quot; x $distx, &amp;quot;\n&amp;quot;; # print a starting line&lt;br /&gt;
&amp;amp;walk($distx, $disty, \$num_red); # pass reference not value&lt;br /&gt;
exit;&lt;br /&gt;
&lt;br /&gt;
sub walk {&lt;br /&gt;
    my ($x, $y, $ref) = @_;&lt;br /&gt;
    my $ct = $$ref;&lt;br /&gt;
    my $prob_red;&lt;br /&gt;
    if ($x == 1 &amp;amp;&amp;amp; $y == 1) { # Reached destination&lt;br /&gt;
	print &amp;quot;*\n&amp;quot;;&lt;br /&gt;
	print &amp;quot;-&amp;quot; x $distx, &amp;quot;\n&amp;quot;; # print the ending line&lt;br /&gt;
	print &amp;quot;reached Belfer after waiting for &amp;quot;, $ct, &amp;quot; red lights\n&amp;quot;;&lt;br /&gt;
    }&lt;br /&gt;
&lt;br /&gt;
    if ($x == 1 &amp;amp;&amp;amp; $y &amp;gt; 1) { # Reached right-side end&lt;br /&gt;
	$prob_red = rand();&lt;br /&gt;
	if ($prob_red &amp;lt; 0.5) { # red light, wait&lt;br /&gt;
	    $ct++;&lt;br /&gt;
	    print &amp;quot;w\n&amp;quot;, &amp;quot; &amp;quot; x ($distx-1);&lt;br /&gt;
	} else {&lt;br /&gt;
	    print &amp;quot;|\n&amp;quot;, &amp;quot; &amp;quot; x ($distx-1);&lt;br /&gt;
	}&lt;br /&gt;
	&amp;amp;walk($x, $y-1, \$ct);&lt;br /&gt;
    }&lt;br /&gt;
&lt;br /&gt;
    if ($x &amp;gt; 1 &amp;amp;&amp;amp; $y == 1) { # Reached bottom&lt;br /&gt;
	$prob_red = rand();&lt;br /&gt;
	if ($prob_red &amp;lt; 0.5) { # red light, wait&lt;br /&gt;
	    $ct++;&lt;br /&gt;
	    print &amp;quot;w&amp;quot;;&lt;br /&gt;
	} else { print &amp;quot;-&amp;quot;}&lt;br /&gt;
	&amp;amp;walk($x-1, $y, \$ct);&lt;br /&gt;
    }&lt;br /&gt;
&lt;br /&gt;
    if ($x &amp;gt; 1 &amp;amp;&amp;amp; $y &amp;gt; 1) {&lt;br /&gt;
	my $prob_across = rand(); # prob of walking right with green light&lt;br /&gt;
	if ($prob_across &amp;gt;= 0.5) { # move one block right&lt;br /&gt;
	    print &amp;quot;-&amp;quot;;&lt;br /&gt;
	    &amp;amp;walk($x-1, $y, \$ct);&lt;br /&gt;
	} else { # red light, move one block down&lt;br /&gt;
	    print &amp;quot;|\n&amp;quot;, &amp;quot; &amp;quot; x ($distx-$x);&lt;br /&gt;
	    &amp;amp;walk($x, $y-1, \$ct);&lt;br /&gt;
	}&lt;br /&gt;
    }&lt;br /&gt;
}&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Jeff&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;python&amp;quot;&amp;gt;&lt;br /&gt;
# In[26]:&lt;br /&gt;
def walk (walker):&lt;br /&gt;
    import random&lt;br /&gt;
    &lt;br /&gt;
    def just_walk (x):                                  #[1]&lt;br /&gt;
        if random.choice([&amp;quot;red_m_direction&amp;quot;, &amp;quot;green_m&amp;quot;])==&amp;quot;green_m&amp;quot;: x[&amp;quot;loca&amp;quot;][&amp;quot;m&amp;quot;] -=1&lt;br /&gt;
        else: x[&amp;quot;loca&amp;quot;][&amp;quot;n&amp;quot;] -=1&lt;br /&gt;
        return x    &lt;br /&gt;
    def may_have_to_wait (dist,waited):                 #[2]&lt;br /&gt;
        if random.choice([&amp;quot;green&amp;quot;,&amp;quot;red&amp;quot;])==&amp;quot;red&amp;quot;: waited += 1&lt;br /&gt;
        else: dist -= 1&lt;br /&gt;
        return (dist,waited)    &lt;br /&gt;
&lt;br /&gt;
    while (0 not in walker[&amp;quot;loca&amp;quot;].values()):         #start walking&lt;br /&gt;
        walker = just_walk(walker) &lt;br /&gt;
    while (walker[&amp;quot;loca&amp;quot;][&amp;quot;m&amp;quot;] !=0):       # if n =0 and m != 0&lt;br /&gt;
        (walker[&amp;quot;loca&amp;quot;][&amp;quot;m&amp;quot;], walker[&amp;quot;waited&amp;quot;]) = may_have_to_wait(walker[&amp;quot;loca&amp;quot;][&amp;quot;m&amp;quot;],walker[&amp;quot;waited&amp;quot;]) &lt;br /&gt;
    while (walker[&amp;quot;loca&amp;quot;][&amp;quot;n&amp;quot;] !=0):       # if m =0 and n != 0&lt;br /&gt;
        (walker[&amp;quot;loca&amp;quot;][&amp;quot;n&amp;quot;], walker[&amp;quot;waited&amp;quot;]) = may_have_to_wait(walker[&amp;quot;loca&amp;quot;][&amp;quot;n&amp;quot;],walker[&amp;quot;waited&amp;quot;]) &lt;br /&gt;
&lt;br /&gt;
    return walker[&amp;quot;waited&amp;quot;]&lt;br /&gt;
&lt;br /&gt;
def given_distance(m,n,rep): &lt;br /&gt;
    waited = list()&lt;br /&gt;
    for x in range(rep):&lt;br /&gt;
        walker={&amp;quot;loca&amp;quot;:{&amp;quot;m&amp;quot;:m,&amp;quot;n&amp;quot;:n}, &amp;quot;waited&amp;quot;:0}     #walker variable&lt;br /&gt;
        waited.append(walk(walker))      &lt;br /&gt;
    return sum(waited)/len(waited)&lt;br /&gt;
# main&lt;br /&gt;
result=list()&lt;br /&gt;
for d in range(1000):                                  # set walking distance here&lt;br /&gt;
        result.append(given_distance (d,d,100))      # set rep here&lt;br /&gt;
import matplotlib.pyplot as plt                     # ploting &lt;br /&gt;
plt.plot(result); plt.show()  &lt;br /&gt;
&lt;br /&gt;
#[1] no reason to wait for ANY red before one of the distances (m or n) is exhausted. Assuming when light is green for m, it must be red for n, vise versa. &lt;br /&gt;
#[2] when one of the directions (m or n = 0), waiting cannot be avoided, start counting. &lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== March 3, 2017 &amp;quot;PiE&amp;quot; (Due 3/10/2017)==&lt;br /&gt;
* Source: Paul Nahin (2008). &amp;quot;Dueling Idiots&amp;quot;, Problem 5.&lt;br /&gt;
* Challenge: obtain numerical values of Pi and E by simulations&lt;br /&gt;
* Simulate pi by [[File:Pi-sim.png|thumbnail]]&lt;br /&gt;
# Randomly generate 10,000 pairs of uniformly distributed numbers from 0 and 1 (simulating throwing darts onto the unit square shown at right)&lt;br /&gt;
# Count the number of points enclosed within the quarter-circle&lt;br /&gt;
# Calculate the pi value from this proportion&lt;br /&gt;
* Simulate e by&lt;br /&gt;
# Generate N random numbers from 0 to 1&lt;br /&gt;
# Divide into N equal-width bins between 0 and 1&lt;br /&gt;
# Count the number of bins Z that receive none of the random numbers&lt;br /&gt;
# Obtain e ~ N/Z (based on binomial sampling formula)&lt;br /&gt;
# Simulate N=1e2, 1e3, and 1e4&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
| &lt;br /&gt;
* By Weigang&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
# simulate pi:&lt;br /&gt;
darts &amp;lt;- sapply(1:1000, function(x) { coords&amp;lt;- runif(2); return(ifelse(coords[1]^2+coords[2]^2 &amp;lt;= 1, 1,0)) })&lt;br /&gt;
pi &amp;lt;- 4*mean(darts)&lt;br /&gt;
&lt;br /&gt;
# simulate e&lt;br /&gt;
N &amp;lt;- 1e3;&lt;br /&gt;
n &amp;lt;- runif(N);&lt;br /&gt;
cts &amp;lt;- numeric();&lt;br /&gt;
for(i in 1:(N-1)) {&lt;br /&gt;
  left &amp;lt;- i/N;&lt;br /&gt;
  right &amp;lt;- (i+1)/N;&lt;br /&gt;
  cts[i] &amp;lt;- length(which(n&amp;gt;=left &amp;amp; n &amp;lt; right))&lt;br /&gt;
}&lt;br /&gt;
e &amp;lt;- N/length(which(cts==0))&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Nicolette&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
#Simulate pi&lt;br /&gt;
random &amp;lt;- 0; square &amp;lt;- 0;&lt;br /&gt;
for (i in 1:1000){&lt;br /&gt;
  random[[i]]&amp;lt;- runif(1,0,1)&lt;br /&gt;
square[[i]]&amp;lt;- sqrt(1-(random[i])^2)&lt;br /&gt;
}&lt;br /&gt;
plot(random, square)&lt;br /&gt;
areaofqc &amp;lt;- (pi/4)&lt;br /&gt;
ranleqc &amp;lt;- length(which(random &amp;lt;=areaofqc))&lt;br /&gt;
squleqc &amp;lt;- length(which(square&amp;lt;=areaofqc))&lt;br /&gt;
&lt;br /&gt;
#Simulate e&lt;br /&gt;
e &amp;lt;- 0;&lt;br /&gt;
for(N in 1:1000) {&lt;br /&gt;
  numbers&amp;lt;-runif(N)&lt;br /&gt;
bin.size&amp;lt;-1/N&lt;br /&gt;
  non.empty&amp;lt;-as.integer(numbers/bin.size)&lt;br /&gt;
  z.empt&amp;lt;- N - length(table(non.empty))&lt;br /&gt;
  e&amp;lt;-c(e, N/z.empt)&lt;br /&gt;
  }&lt;br /&gt;
plot(e)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By Brian&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
#When you throw a dart at the a unit square dartboard the Probability of hitting the portion of the circle radius with center=(0,0) inside the unit square is  pi/4.  We can say that hitting the 1/4 circle within the unit square with a dart is a bernoulli random variable where p=pi/4. Further, we can say E(bernoulli=hit)=pi/4. &lt;br /&gt;
# Imagine you are scoring the game of darts as follows- 1 point if you throw in the 1/4 circle and 0 points if you miss this space. &lt;br /&gt;
# If you throw 10000 darts you just did 10000 iid bernoulli trials. By the law of large numbers if we count the number of hits to the 1/4 circle of radius 1 and divide by number of darts thrown we will get something pretty close to E(beroulli) . Multiply that number by 4 and you have an estimate of Pi.  &lt;br /&gt;
&lt;br /&gt;
## Generate 10,000 pairs of points in the unit square with uniform distribution&lt;br /&gt;
&lt;br /&gt;
x&amp;lt;-c(runif(1000000, min = 0, max = 1))&lt;br /&gt;
y&amp;lt;-c(runif(1000000, min = 0, max = 1))&lt;br /&gt;
point&amp;lt;-data.frame(x,y)&lt;br /&gt;
&lt;br /&gt;
plot(point$x,point$y)&lt;br /&gt;
point.sub&amp;lt;-subset(point,y&amp;lt;=sqrt(1-x^2))&lt;br /&gt;
&lt;br /&gt;
plot(point.sub$x,point.sub$y)&lt;br /&gt;
&lt;br /&gt;
z&amp;lt;-4*nrow(point.sub)/1000000&lt;br /&gt;
z&lt;br /&gt;
error&amp;lt;-pi-z&lt;br /&gt;
error&lt;br /&gt;
&lt;br /&gt;
## simulating exp&lt;br /&gt;
#It&#039;s a well that a binomial with large n and tiny p is a good estimate of the poisson random For this estimate lambda=np.  In this case n=10,000 and p=the probability of falling into a particular unit=1/10000....... &lt;br /&gt;
 simexp&amp;lt;-function (n) {&lt;br /&gt;
a&amp;lt;-c(runif(n,min=0,max=1))&lt;br /&gt;
b&amp;lt;-c()&lt;br /&gt;
&lt;br /&gt;
for (i in 1:n) {&lt;br /&gt;
  c&amp;lt;-subset(a,(i-1)/n&amp;lt;a &amp;amp; a&amp;lt;=i/n)&lt;br /&gt;
  d&amp;lt;-length(c)&lt;br /&gt;
  b&amp;lt;-append(b,d)&lt;br /&gt;
}&lt;br /&gt;
bzero&amp;lt;-subset(b,b==0)&lt;br /&gt;
length(bzero)&lt;br /&gt;
print(n/length(bzero))&lt;br /&gt;
}&lt;br /&gt;
simexp(1000)&lt;br /&gt;
simexp(10000)&lt;br /&gt;
simexp(100000)&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*By John&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
#python code&lt;br /&gt;
# Simulate pi&lt;br /&gt;
simulation = 10000&lt;br /&gt;
distances = np.array([pow(uniform(0, 1), 2) + pow(uniform(0, 1), 2) for i in range(simulation)])&lt;br /&gt;
pi = count_nonzero(distances &amp;lt; 1) / simulation * 4&lt;br /&gt;
pi&lt;br /&gt;
&lt;br /&gt;
# Simulate e&lt;br /&gt;
total = 10000&lt;br /&gt;
ranges = [[x/total, (x+1)/total] for x in range(total)]&lt;br /&gt;
sample_space = [uniform(0, 1) for x in range(total)]&lt;br /&gt;
index = []&lt;br /&gt;
for i in range(total):&lt;br /&gt;
    if any(ranges[i][0] &amp;lt;= point &amp;lt;= ranges[i][1] for point in sample_space):&lt;br /&gt;
        index.append(i)&lt;br /&gt;
e = total / (total - len(set(index)))&lt;br /&gt;
e&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Feb 24, 2017 &amp;quot;Idiots&amp;quot; (Due 3/3/2017)==&lt;br /&gt;
[[File:Idiots.png|thumbnail]]&lt;br /&gt;
* Source: Paul Nahin (2000), &amp;quot;Dueling Idiots&amp;quot;. Problem #2: &amp;quot;When Idiots Duel&amp;quot;&lt;br /&gt;
* Game: Idiots A and B decide to duel with a gun in the following way: They will insert a single bullet into the gun&#039;s cylinder. The cylinder has a total of 6 slots. Idiot A will spin the cylinder and shoot at B. If the gun doesn&#039;t fire, then A will give the gun to B, who will spin the cylinder and then shoot at A. This back-and-forth duel will continue until one fool shoots (and kills) the other. &lt;br /&gt;
* Questions: (1) what is the probability that A will win (and B dies); (2) What is the average number of trigger pulls before someone dies?&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
| &lt;br /&gt;
* By Roy&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
gun &amp;lt;-c(1,0,0,0,0,0)&lt;br /&gt;
gun2 &amp;lt;- c(1,0,0,0,0,0)&lt;br /&gt;
deadman &amp;lt;-0 ;idiot1win &amp;lt;-0; idiot2win &amp;lt;- 0; nooned &amp;lt;- 0; total &amp;lt;- 0&lt;br /&gt;
 while(deadman &amp;lt; 1000){&lt;br /&gt;
  idiot1shot &amp;lt;-sample(gun)&lt;br /&gt;
  if (length(which(idiot1shot[1] == 1))){&lt;br /&gt;
    deadman &amp;lt;- deadman + 1&lt;br /&gt;
    idiot1win &amp;lt;- idiot1win + 1&lt;br /&gt;
  }else{&lt;br /&gt;
      idiot2shot &amp;lt;-sample(gun2)&lt;br /&gt;
      if (length(which(idiot2shot[1] == 1))){&lt;br /&gt;
        deadman &amp;lt;- deadman +1&lt;br /&gt;
        idiot2win &amp;lt;- idiot2win +1&lt;br /&gt;
      } else {nooned &amp;lt;- nooned + 1}&lt;br /&gt;
  }&lt;br /&gt;
  total &amp;lt;- total + 1&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
deadman&lt;br /&gt;
idiot1win&lt;br /&gt;
idiot2win&lt;br /&gt;
nooned&lt;br /&gt;
total&lt;br /&gt;
&lt;br /&gt;
p &amp;lt;- idiot1win/1000&lt;br /&gt;
p*100&lt;br /&gt;
takes2kill &amp;lt;- deadman/total&lt;br /&gt;
takes2kill*100&lt;br /&gt;
idiot1shot&lt;br /&gt;
idiot2shot&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* By Weigang&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
rounds &amp;lt;- sapply(1:1000, function(x) {&lt;br /&gt;
  alive &amp;lt;- 1;&lt;br /&gt;
  round &amp;lt;- 0;&lt;br /&gt;
  while(alive == 1){&lt;br /&gt;
    spin &amp;lt;- sample(c(0,0,0,0,0,1));&lt;br /&gt;
    round &amp;lt;- round + 1;&lt;br /&gt;
    if(spin[1] == 1) { alive &amp;lt;- 0 }&lt;br /&gt;
  }&lt;br /&gt;
  return(round)&lt;br /&gt;
})&lt;br /&gt;
prob.a.live &amp;lt;- length(which(rounds %% 2 == 0))&lt;br /&gt;
barplot(table(rounds)/1000, xlab=&amp;quot;num rounds&amp;quot;, ylab=&amp;quot;Prob&amp;quot;, main = &amp;quot;Sudden Death with a 6-slot gun (sim=1000)&amp;quot;, las=1)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Feb 17, 2017 &amp;quot;Birthday&amp;quot; (Due 2/24/2017)==&lt;br /&gt;
[[File:B-day.png|thumbnail]]&lt;br /&gt;
* Problem: What is the probability NONE of the N people in a room sharing a birthday?&lt;br /&gt;
# Randomly select N individuals and record their B-days&lt;br /&gt;
# Count the B-days NOT shared by ANY two individuals&lt;br /&gt;
# Repeat (for each N) 1000 times, and obtain probability by averaging the previous counts (i.e., divided by 1000)&lt;br /&gt;
# Vary N from 10 to 100, increment by 10&lt;br /&gt;
# Plot probability of no-shared B-Day (Y-axis) versus N (x-axis), with either a stripchart or boxplot, or both&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
| &lt;br /&gt;
* By Roy&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
N &amp;lt;- 0&lt;br /&gt;
samples &amp;lt;- 0;&lt;br /&gt;
prob.nodub &amp;lt;-0;&lt;br /&gt;
for(j in 10:100){&lt;br /&gt;
  counting.no.dups &amp;lt;- 0;&lt;br /&gt;
  test &amp;lt;- for(i in 1:1000){&lt;br /&gt;
  bdays &amp;lt;- sample(seq(as.Date(&#039;1990/01/01&#039;), as.Date(&#039;1990/12/31&#039;), by=&amp;quot;day&amp;quot;), N, replace=T)&lt;br /&gt;
  dups &amp;lt;- duplicated(bdays, incomparables = FALSE)&lt;br /&gt;
  ch &amp;lt;-length(which(dups == TRUE))&lt;br /&gt;
    if(ch==0){&lt;br /&gt;
      counting.no.dups &amp;lt;- counting.no.dups +1&lt;br /&gt;
    }&lt;br /&gt;
    fine &amp;lt;- (counting.no.dups/1000)&lt;br /&gt;
  }&lt;br /&gt;
  print(N)&lt;br /&gt;
  print(fine)&lt;br /&gt;
  N &amp;lt;- N + 1&lt;br /&gt;
  samples[[j]] &amp;lt;- N&lt;br /&gt;
  prob.nodub[[j]] &amp;lt;- fine&lt;br /&gt;
}&lt;br /&gt;
plot(samples,prob.nodub, main=&amp;quot;Birthday Simulation&amp;quot;, xlab = &amp;quot;Sample Size&amp;quot;, ylab = &amp;quot;Probability of no Duplicates&amp;quot;, las =1)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* By weigang&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
days &amp;lt;- 1:365;&lt;br /&gt;
&lt;br /&gt;
find.overlap &amp;lt;- function(x) { return(length(which(table(x)&amp;gt;1))) }&lt;br /&gt;
&lt;br /&gt;
output &amp;lt;- sapply(1:100, function(x) { # num of people in the room&lt;br /&gt;
  ct.no.overlap &amp;lt;- 0;&lt;br /&gt;
  for (k in 1:100) {&lt;br /&gt;
    bdays &amp;lt;- sample(days, x, replace = T);&lt;br /&gt;
    ct &amp;lt;- find.overlap(bdays);&lt;br /&gt;
    if (!ct) { ct.no.overlap &amp;lt;- ct.no.overlap + 1}&lt;br /&gt;
  }&lt;br /&gt;
  return(ct.no.overlap);&lt;br /&gt;
})&lt;br /&gt;
 &lt;br /&gt;
plot(1:100, output/100, xlab=&amp;quot;Group size&amp;quot;, ylab = &amp;quot;Prob (no shared b-day)&amp;quot;, las=1, main=&amp;quot;B-day (sim=100 times)&amp;quot;, type=&amp;quot;l&amp;quot;)&lt;br /&gt;
abline(h=0.5, lwd=2, col=2, lty=2)&lt;br /&gt;
abline(v=seq(0, 100, 5), col=&amp;quot;gray&amp;quot;)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Feb 10, 2017 &amp;quot;Dating&amp;quot; (Valentine&#039;s Day Special; Due 2/17/2017)==&lt;br /&gt;
[[File:Dating.png|thumbnail]]&lt;br /&gt;
* Source: Paul Nahin (2008), &amp;quot;Digital Dice&amp;quot;. Problem #20: &amp;quot;An Optimal Stopping Problem&amp;quot;&lt;br /&gt;
* Problem: What is the optimal time point when one should stop dating more people and settle on a mate choice (and live with the decision)&lt;br /&gt;
* Your best strategy is to date an initial sample of N individuals, rejecting all, and marry the next one ranked higher than any of your N individuals. The question is what is the optimal number for N.&lt;br /&gt;
# The problem could be investigated by simulating a pool of 10 individuals, ranked from 1-10 (most desirable being 1) and then take a sample of N&lt;br /&gt;
# You may only date one individual at a time&lt;br /&gt;
# You cannot go back to reach previously rejected candidates&lt;br /&gt;
# Simulate N from 0 to 9 (zero means marrying the first date, a sample size of zero)&lt;br /&gt;
# For each N, obtain the probability of finding the perfect mate (i.e., ranked 1st) by running simulation 1000 times&lt;br /&gt;
# Plot barplot of probability versus sample size N.&lt;br /&gt;
# Expected answer: N=4&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
| &lt;br /&gt;
* By Weigang&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
pick.candidate &amp;lt;- function(min, array) {&lt;br /&gt;
  for (i in 1:length(array)) {&lt;br /&gt;
    if (array[i] &amp;lt; min) {&lt;br /&gt;
      return(array[i])&lt;br /&gt;
    } else {next}&lt;br /&gt;
  }&lt;br /&gt;
  return(0) # No 1. has been sampled and rejected&lt;br /&gt;
}&lt;br /&gt;
candidates &amp;lt;- 1:10;&lt;br /&gt;
output &amp;lt;- sapply(0:9, function(x) {&lt;br /&gt;
  ct &amp;lt;- 0;&lt;br /&gt;
  for(k in 1:1000) {&lt;br /&gt;
    if (x==0) { # no sample, marry the 1st guy &lt;br /&gt;
      sampled &amp;lt;- sample(candidates, 1);&lt;br /&gt;
      if (sampled == 1) {ct &amp;lt;- ct+1}&lt;br /&gt;
    } else {&lt;br /&gt;
      sampled &amp;lt;- sample(candidates, x);&lt;br /&gt;
      not.sampled &amp;lt;- candidates[-sampled];&lt;br /&gt;
      not.sampled &amp;lt;- sample(not.sampled);&lt;br /&gt;
      if (pick.candidate(min =  min(sampled), array = not.sampled) == 1) {ct &amp;lt;- ct+1}&lt;br /&gt;
    }&lt;br /&gt;
  }&lt;br /&gt;
  return(ct);&lt;br /&gt;
})&lt;br /&gt;
barplot(output/1e3, names.arg = 0:9, xlab = &amp;quot;number of sampled dates&amp;quot;, las=1, main = &amp;quot;Optimal stopping for dating (N=10 candidates)&amp;quot;, ylab = &amp;quot;Prob(marrying No.1)&amp;quot;)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Feb 3, 2017 &amp;quot;US Presidents&amp;quot; (Due 2/10/2017)==&lt;br /&gt;
[[File:Sim-presidents.png|thumbnail]]&lt;br /&gt;
* Download [[File:Presidents.txt|thumbnail]]: 1st column is the order, 2nd column is the name, the 3rd column is the year of inauguration; tab-separated&lt;br /&gt;
* Your job is to create an R, Perl, or Python script called “us-presidents”, which will&lt;br /&gt;
# Read the table&lt;br /&gt;
# Store the original/correct order&lt;br /&gt;
# Shuffle/permute the rows and record the new order&lt;br /&gt;
# Count the number of matching orders&lt;br /&gt;
# Repeat Steps 3-4 for a 1000 times&lt;br /&gt;
# Plot histogram or barplot (better) to show distribution of matching counts&lt;br /&gt;
# Hint: For R, use the sample() function. For Perl, use the rand() function.&lt;br /&gt;
{| class=&amp;quot;wikitable sortable mw-collapsible&amp;quot;&lt;br /&gt;
|- style=&amp;quot;background-color:lightsteelblue;&amp;quot;&lt;br /&gt;
! scope=&amp;quot;col&amp;quot; style=&amp;quot;width: 300px;&amp;quot; | Submitted Codes&lt;br /&gt;
|- style=&amp;quot;background-color:powderblue;&amp;quot;&lt;br /&gt;
| &lt;br /&gt;
* By Mei&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
pres.list &amp;lt;- lapply(1:1000, function(x) pres[sample(nrow(pres)),])&lt;br /&gt;
cts &amp;lt;- sapply(pres.list, function(x) {&lt;br /&gt;
  ct.match &amp;lt;- 0;&lt;br /&gt;
  for (i in 1:45){&lt;br /&gt;
    if (pres$order[i] == x[i,1]){&lt;br /&gt;
      #cat(as.character(x[i,2]), x[i,1], &amp;quot;\n&amp;quot;)  #as.character to avoid the factor info&lt;br /&gt;
      ct.match &amp;lt;- ct.match + 1;&lt;br /&gt;
      #cat(ct.match)&lt;br /&gt;
    }&lt;br /&gt;
  }&lt;br /&gt;
  ct.match #return&lt;br /&gt;
})&lt;br /&gt;
barplot(table(cts),xlab = &amp;quot;Number of matches per shuffle&amp;quot;, border = &amp;quot;hotpink&amp;quot;, col = &amp;quot;pink&amp;quot;, ylab = paste(&amp;quot;Frequency total of:&amp;quot;,length(pres.list)), main = &amp;quot;US Presidents&amp;quot;)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
* By John&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;python&amp;quot;&amp;gt;&lt;br /&gt;
import pandas as pd&lt;br /&gt;
import numpy as np&lt;br /&gt;
from pandas import DataFrame&lt;br /&gt;
import matplotlib.pyplot as plt&lt;br /&gt;
&lt;br /&gt;
df = pd.read_table(&amp;quot;presidents.txt&amp;quot;, names=[&amp;quot;num&amp;quot;, &amp;quot;name&amp;quot;, &amp;quot;presidency&amp;quot;])&lt;br /&gt;
&lt;br /&gt;
name_list = list(df.name)&lt;br /&gt;
&lt;br /&gt;
# create a list to store matches after each shuffle&lt;br /&gt;
shuffle_record = []&lt;br /&gt;
&lt;br /&gt;
# create a function in which the first argument is the original dataset; second argument is number of shuffles&lt;br /&gt;
def shuffler2(original, n):&lt;br /&gt;
    record = []&lt;br /&gt;
    for i in range(n):&lt;br /&gt;
        num = 0&lt;br /&gt;
        each_shuffle = {}&lt;br /&gt;
        temp = original.reindex(np.random.permutation(original.index)) # do shuffling for each&lt;br /&gt;
        compare = original.num == temp.num&lt;br /&gt;
        matched_df = original.ix[compare[compare == True].index]&lt;br /&gt;
        for i in matched_df.index:&lt;br /&gt;
            each_shuffle[i] = matched_df.name[i]&lt;br /&gt;
        shuffle_record.append(each_shuffle)&lt;br /&gt;
        try:&lt;br /&gt;
            num = compare.value_counts()[1]&lt;br /&gt;
        except:&lt;br /&gt;
            pass&lt;br /&gt;
        record.append(num)&lt;br /&gt;
    return(record)&lt;br /&gt;
result2 = shuffler2(df, 1000)&lt;br /&gt;
&lt;br /&gt;
plt.hist(result2, color=&amp;quot;yellow&amp;quot;)&lt;br /&gt;
plt.title(&amp;quot;Histgram for President Data&amp;quot;)&lt;br /&gt;
plt.show()&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt; &lt;br /&gt;
* By Weigang&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
p &amp;lt;- read.table(&amp;quot;Presidents.txt&amp;quot;, sep=&amp;quot;\t&amp;quot;, header=F)&lt;br /&gt;
colnames(p) &amp;lt;- c(&amp;quot;order&amp;quot;, &amp;quot;name&amp;quot;, &amp;quot;inaug.year&amp;quot;)&lt;br /&gt;
# Use &amp;quot;sapply&amp;quot; or &amp;quot;lapply&amp;quot; for loops: no need to pre-define a vector to store results&lt;br /&gt;
p.sim &amp;lt;- sapply(1:10000, function (x) {&lt;br /&gt;
  length(which(sample(p$order) == p$order))&lt;br /&gt;
  }) &lt;br /&gt;
barplot(table(p.sim)/1e4, las=1)&lt;br /&gt;
p.exp &amp;lt;- rpois(1e4, 1) # draw 10000 Poisson random deviates&lt;br /&gt;
mp &amp;lt;- barplot(table(p.exp)/1e4, las=1, xlab = &amp;quot;Num of matching presidents&amp;quot;) # mid-point on x-axis&lt;br /&gt;
lines(mp[1:7], table(p.sim)/1e4, type=&amp;quot;b&amp;quot;, col=2) # add a line (Poisson-expected) to the barplot (simulated)&lt;br /&gt;
legend(&amp;quot;topright&amp;quot;, c(&amp;quot;Simulated&amp;quot;, &amp;quot;Poisson expectation&amp;quot;), col=1:2, lty=1)&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
|}&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Main_Page&amp;diff=6705</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Main_Page&amp;diff=6705"/>
		<updated>2026-07-02T10:48:43Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Course/Lecture syllabi */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{|&lt;br /&gt;
|-&lt;br /&gt;
| &lt;br /&gt;
[[File:Protrait-July-2024-Qiu.jpg| x300 px | thumb |&lt;br /&gt;
&amp;lt;strong&amp;gt;Welcome to Qiu Lab Wiki @ Hunter&amp;lt;/strong&amp;gt;&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;Weigang Qiu, Ph.D., Professor, [https://biology.hunter.cuny.edu Department of Biological Sciences]&amp;lt;br&amp;gt; &lt;br /&gt;
[https://hunter.cuny.edu Hunter College] of [https://www.cuny.edu City University of New York]&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
Adjunct Faculty, Weill Cornell Medical College, [https://weill.cornell.edu/units/systems-and-computational-biomedicine Department of Systems and Computational Biomedicine]&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
Belfer Research Building, Room 402&amp;lt;br&amp;gt;413 East  69th Street, New York, NY 10021&amp;lt;br&amp;gt;Office: 1-212-896-0445&amp;lt;br&amp;gt;Email: wqiu-at-(hunter.cuny.edu) &amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
[https://goo.gl/maps/xv1KmaW3XEnxaY1V7 Directions by Google Map]]]&lt;br /&gt;
|&lt;br /&gt;
[[File:Book-cover.jpg | x300 px | thumb | A textbook for microbial genomics &amp;amp; evolution (Spring 2026):&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;[https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 Read excerpts &amp;amp; Order from the publisher using Author Discount (25%) Code &amp;quot;C2508&amp;quot;]&amp;lt;br&amp;gt;&amp;lt;br&amp;gt; [https://digitalgenomes.net Companion codebook (work in progress &amp;amp;#x1F60A;)]]]&lt;br /&gt;
|&lt;br /&gt;
__TOC__&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Fieldwork Gallery==&lt;br /&gt;
&amp;lt;gallery heights=&amp;quot;200&amp;quot; perrow=&amp;quot;4&amp;quot; style=&amp;quot;text-align:left&amp;quot; widths=&amp;quot;250&amp;quot; mode=&amp;quot;packed&amp;quot;&amp;gt;&lt;br /&gt;
File:Tick trip rockefeller.jpeg|Apr 2026, Rockefeller State Park, NY&lt;br /&gt;
File:Tick-trip-2023.jpg|Oct 2023, Long Island, NY&lt;br /&gt;
File:Tick-trip-2022.jpg|June 2022, MA&lt;br /&gt;
File:Tick-trip-2021.jpg|June 2021, MA&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&amp;lt;youtube width=&amp;quot;400&amp;quot; height=&amp;quot;300&amp;quot;&amp;gt;Jb4ACK-GjM0&amp;lt;/youtube&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Lab publications==&lt;br /&gt;
===Lyme Genomics, Evolution, &amp;amp; Ecology===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;4&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
&lt;br /&gt;
File:Screenshot_2024-08-15_094908.png | link=https://journals.asm.org/doi/10.1128/mbio.01749-24 | Akther et al. 2024. &amp;quot;Natural selection and recombination at host-interacting lipoprotein loci drive genome diversification of Lyme disease and related bacteria&amp;quot; &#039;&#039;&#039;&#039;&#039;mBio&#039;&#039;&#039;&#039;&#039; 0:e01749-24. Press coverage: [https://www.gc.cuny.edu/news/cuny-graduate-center-biologists-map-dna-lyme-disease-bacteria CUNY Graduate Center News Story]; [https://hunter.cuny.edu/news/hunter-researcher-maps-dna-of-lyme-disease-bacteria/ Hunter press]&lt;br /&gt;
&lt;br /&gt;
File:Applsci-13-11587-g004.png | link=https://doi.org/10.3390/app132011587 | Di, Chong, Brian Sulkow, Weigang Qiu, and Shipeng Sun. 2023. &amp;quot;Effects of Micro-Scale Environmental Factors on the Quantity of Questing Black-Legged Ticks in Suburban New York&amp;quot; &#039;&#039;&#039;&#039;&#039;Applied Sciences&#039;&#039;&#039;&#039;&#039; 13, no. 20: 11587.&lt;br /&gt;
&lt;br /&gt;
File:Spectrum.01743-22-f003.gif | link=https://doi.org/10.1128/spectrum.01743-22  | Li, Di, Zeglis, Qiu (2022). “Evolution of the &#039;&#039;vls&#039;&#039; antigenic variability locus of the Lyme Disease pathogen and development of recombinant monoclonal antibodies targeting conserved VlsE epitopes”. &#039;&#039;&#039;&#039;&#039;Microbial Spectrum&#039;&#039;&#039;&#039;&#039; 10 (5):1-15. &lt;br /&gt;
&lt;br /&gt;
File:Fig7-small.png | link= https://pubmed.ncbi.nlm.nih.gov/34413477 | Di*, Akther*, Bezrucenkovas, Ivanova, Sulkow, Wu, Mneimneh, Gomes-Solecki, Qiu (2021). &amp;quot;Maximum antigen diversification in a lyme bacterial population and evolutionary strategies to overcome pathogen diversity&amp;quot;. &#039;&#039;&#039;&#039;&#039;The ISME Journal&#039;&#039;&#039;&#039;&#039;. 16, 447-464. (*co-first authors) [https://communities.springernature.com/posts/jenner-s-dilemma-and-how-to-win-evolutionary-arms-races-against-microbial-pathogens Blog Post]&lt;br /&gt;
&lt;br /&gt;
File:Ira-fig2.png | link=https://doi.org/10.21775/9781913652616 | Schwartz, Margos, Casjens, Qiu, Eggers (2020). &amp;quot;Multipartite Genome of Lyme Disease &#039;&#039;Borrelia&#039;&#039;: Structure, Variation and Prophages&amp;quot;. &#039;&#039;&#039;&#039;&#039;Current Issues in Molecular Biology&#039;&#039;&#039;&#039;&#039;. 42:409-454. &lt;br /&gt;
&lt;br /&gt;
File:Fig6-1-Barbour.png | link=https://doi.org/10.1002/9781118960608.gbm01525 | Barbour &amp;amp; Qiu (2019). &#039;&#039;Borreliella&#039;&#039;. In &#039;&#039;&#039;&#039;&#039;Bergey&#039;s Manual of Systematics of Archaea and Bacteria&#039;&#039;&#039;&#039;&#039;. John Wiley &amp;amp; Sons, Inc., in association with Bergey&#039;s Manual Trust. &lt;br /&gt;
&lt;br /&gt;
File:Zjm9990961420002.jpeg | link=https://doi.org/10.1128/jcm.00940-18 | Di L, Wan Z, Akther S, Ying CX, Larracuente A, Li L, Di C, Nunez R, Cucura DM, Goddard NL, Krampis K, Qiu WG. (2018). Genotyping and Quantifying Lyme Pathogen Strains by Deep Sequencing of the Outer Surface Protein C (&#039;&#039;ospC&#039;&#039;) Locus. &#039;&#039;&#039;Journal of Clinical Microbiology&#039;&#039;&#039;. 56(11):e00940-18. &lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Evolution &amp;amp; Learning Algorithms===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;3&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File:Pathogens-12-00388-g003.png | link=https://doi.org/10.3390/pathogens12030388 | Ely, Koh, Ho, Hassan, Pham, Qiu (2023). Novelty Search Promotes Antigenic Diversity in Microbial Pathogens. &#039;&#039;&#039;&#039;&#039;Pathogens&#039;&#039;&#039;&#039;&#039; 12:388.&lt;br /&gt;
File:Pone.0224552.g005.png | link=https://doi.org/10.1371/journal.pone.0224552 | Attie, Sulkow, Di, Qiu (2019). Genetic codes optimized as a traveling salesman problem. &#039;&#039;&#039;&#039;&#039;PLoS ONE&#039;&#039;&#039;&#039;&#039; 14(10): e0224552.&lt;br /&gt;
File:Pcbi.1005677.g001.png | link=https://doi.org/10.1371/journal.pcbi.1005677 | Yan, Deforet, Boyle, Rahman, Liang, Okegbe, Dietrich, Qiu, Xavier (2017). Bow-tie signaling in c-di-GMP: Machine learning in a simple biochemical network. &#039;&#039;&#039;&#039;&#039;PLoS Comput Biol&#039;&#039;&#039;&#039;&#039; 13(8): e1005677. &lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Informatics Tool Development===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File: 12859 2018 2074 Fig1 HTML.png | link=https://pubmed.ncbi.nlm.nih.gov/29499649 | Hernandez, Bernstein, Pagan, Vargas, McCaig, Ramrattan, Akther, Larracuente, Di, Vieira, Qiu. (2018). BbWrapper: BioPerl-based sequence and tree utilities for rapid prototyping of bioinformatics pipelines. &#039;&#039;&#039;&#039;&#039;BMC Bioinformatics&#039;&#039;&#039;&#039;&#039; 19(1):76.&lt;br /&gt;
File: 12859 2014 Article 6488 Fig3 HTML.png | link=https://pubmed.ncbi.nlm.nih.gov/24994456 | Di, Pagan, Packer, Martin, Akther, Ramrattan, Mongodin, Fraser, Schutzer, Luft, Casjens and Qiu. (2014). BorreliaBase: a phylogeny-centered browser of &#039;&#039;Borrelia&#039;&#039; genomes. &#039;&#039;&#039;&#039;&#039;BMC Bioinformatics&#039;&#039;&#039;&#039;&#039; 15:233.&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*[http://scholar.google.com/citations?hl=en&amp;amp;user=Ds6u39QAAAAJ) Full list by Google Scholar]&lt;br /&gt;
*[http://www.ncbi.nlm.nih.gov/sites/myncbi/weigang.qiu.1/bibliography/42770924/public/ Full list by NCBI Bibliography]&lt;br /&gt;
last update: March 20, 2023&lt;br /&gt;
&lt;br /&gt;
==Lab members and trainees==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
!Year/Period !!Doctoral members &amp;amp; trainees !!Other members &amp;amp; trainees&lt;br /&gt;
|-&lt;br /&gt;
|Current Academic Year&lt;br /&gt;
(Spring &amp;amp; Summer 2026)&lt;br /&gt;
||&lt;br /&gt;
*Edee Amin: CUNY Graduate Center, Bio/EEB doctoral program&lt;br /&gt;
* Dr Yozen Hernandez: System administrator (part-time), Ph.D. from Boston University&lt;br /&gt;
||&lt;br /&gt;
*Tasmina Hassan: Hunter Bio/CS&lt;br /&gt;
*Mathew DiCicco: Hunter Math/CS&lt;br /&gt;
*Tara Doma Lama: Hunter Bio/Bioinformatics&lt;br /&gt;
*Hagar Abuzaid: Hunter Bio&lt;br /&gt;
|-&lt;br /&gt;
|Alumni (Since Fall 2002)&lt;br /&gt;
||&lt;br /&gt;
* Brandon Ely (2026): CUNY Grad Center, Biology/MCD doctoral program&lt;br /&gt;
* Li Li (Lily, 2023): CUNY Grad Center, Biology/EEB doctoral program&lt;br /&gt;
*Dr Lia Di: Ph.D. Postdoctoral Research Associate, from Nanjing Agricultural University &amp;amp; Wisconsin Blood Institute&lt;br /&gt;
*Dr Saymon Akther (2022): CUNY Grad Center, Biology/EEB&lt;br /&gt;
*Dr Rayees Rahman: Hunter Bio/Bioinformatics, Ph.D. from Mt Sinai Medical School&lt;br /&gt;
*Dr Che Martin (2013): CUNY Grad Center, Biology/MCD&lt;br /&gt;
*Dr James Haven (2011): CUNY Grad Center, Biology/MCD&lt;br /&gt;
*Dr Tika Sukarna (2009): CUNY Grad Center, Biology/MCD&lt;br /&gt;
*Dr Juan Coronado (2008): CUNY Grad Center, Biology/MCD (Dr Peter Lipke)&lt;br /&gt;
*Dr William McCaig: CUNY BA, Ph.D. from Stony Brook University&lt;br /&gt;
*Dr Vincent Xue: CUNY CS/Bioinformatics, Ph.D. from MIT&lt;br /&gt;
*Dr Fubin Li: CUNY Grad Center, Biology/MCD (Dr Laurel Eckhardt)&lt;br /&gt;
*Dr Oliver Attie: Postdoctoral Research Associate, Ph.D. from NYU&lt;br /&gt;
||&lt;br /&gt;
(published coauthors)&amp;lt;br&amp;gt; &lt;br /&gt;
*Brian Sulkow: CUNY Grad Center, Mathematics&lt;br /&gt;
* Winston Koh: Hunter Bio/CS&lt;br /&gt;
*Eamen Ho: Hunter Bio/Bioinformatics&lt;br /&gt;
*Ahn Pham: Hunter Bio/Bioinformatics&lt;br /&gt;
*Chris Panlasigui: Hunter Bio/Bioinformatics&lt;br /&gt;
*Amanda Amanda Larracuente: CUNY Grad Center, Biology/MCD&lt;br /&gt;
* Pedro Pagan: Hunter Bio/Bioinformatics&lt;br /&gt;
*Edgaras Bezrucenkovas: Hunter Chem/Bioinformatics&lt;br /&gt;
*Girish Ramrattan: Hunter Bio/Bioinformatics&lt;br /&gt;
* Levy Vargas: Hunter Bio/Bioinformatics&lt;br /&gt;
*Chong Di: Hunter Geography&lt;br /&gt;
*Roy Nunez: Hunter Bio&lt;br /&gt;
*Mei Wu: CUNY City Tech&lt;br /&gt;
*Desiree Pante: Hunter Bio&lt;br /&gt;
*Saimtun Shipa: Hunter Stat (MA)&lt;br /&gt;
*Bing Wu: Hunter Bio/Biotechnology&lt;br /&gt;
* Svidatoslav Kendall (Slav): Hunter Biology&lt;br /&gt;
*Philip Romov: Hunter CS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Web Apps==&lt;br /&gt;
===Apps with Collaborators (or from published papers)===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;4&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File:Screenshot_2024-06-19_205621.png | link=https://cov.genometracker.org/finalelisa-app/ | Borrelia diagnostic antigens (App developed by Liann Aris-Henry; Data from [https://journals.asm.org/doi/10.1128/jcm.01142-19 Arumugam et al (2019)])&lt;br /&gt;
File:Screenshot_2024-06-18_144931.png | link=https://cov.genometracker.org/borrelia-app/ | Borrelia growth transcriptome (App developed by Laziz Asamov; Data from [https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0164165 Arnold et al (2016)])&lt;br /&gt;
File: Silac.png | link=http://borreliabase.org/~wgqiu/Polotskaia_etal_2014/supp-fig-s1/ | Data &amp;amp; [http://borreliabase.org/~wgqiu/Polotskaia_etal_2014/supp-table-s1/ GSEA] associated with Polotskaia_etal_2014 (with Bargonetti Lab @Hunter)&lt;br /&gt;
File:Screenshot_2023-03-24_122016.png | link=http://borreliabase.org/~wgqiu/clickme-khalikuz/temp-Points.html | Phosphoproteome MDM2KD data set (with Bargonetti Lab @Hunter)&lt;br /&gt;
File:Screenshot_2023-03-24_122312.png | link=http://borreliabase.org/~wgqiu/mpai-v3/ | Genes &amp;amp; Pathways Associated with mTP53 (with Bargonetti Lab @Hunter)&lt;br /&gt;
File:Spombe.png | link=Spombe | S. pombe transcriptomes (with Zhong Lab @Hunter)&lt;br /&gt;
File:Screenshot 2023-03-24 122223.png | link=http://borreliabase.org/~wgqiu/gbs-browser-v3/ | Group B Streptococcus (GBS) genome browser (with Wu Lab @Shenzhen)&lt;br /&gt;
File:Screenshot 2023-03-24 122245.png | link=http://borreliabase.org/~wgqiu/oneKGenome/ | 1K genome (with Ogunwobi Lab @Hunter)&lt;br /&gt;
File:Screenshot 2023-03-24 122450.png | link=http://borreliabase.org/~wgqiu/E_faecalis/ | E faecalis genome browser (with Morales Lab @WCMC)&lt;br /&gt;
File:Screenshot 2023-03-24 122145.png | link=http://borreliabase.org/~wgqiu/carmen-proteomics/ | Proteomics GSEA results (with Melendez Lab @Hunter) &lt;br /&gt;
File:Screenshot 2023-03-24 122051.png | link=http://borreliabase.org/~wgqiu/Tcell-pHis-ACS/ | pHis Protein Browser (with Skolnik lab @NYU)&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Qiu Lab Apps===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;3&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File:Screenshot 2023-03-24 122558.png | link=http://borreliabase.org | Lyme genome browser&lt;br /&gt;
File:Screenshot 2023-03-24 133320.png | link=http://borreliabase.org/~wgqiu/pf-trees/ | Lyme pathogen plasmid partitioning gene trees&lt;br /&gt;
File:Screenshot 2023-03-24 133636.png | link=http://borreliabase.org/vls-finder/ | vls Finder in Lyme pathogen genomes&lt;br /&gt;
File:Screenshot 2023-03-24 122359.png | link=http://borreliabase.org/~wgqiu/code-wheel | Codon Wheel&lt;br /&gt;
File:PhyloView.png | link=http://borreliabase.org/~wgqiu/PhyloView | Co-visualization of a tree with an alignment and character matrix&lt;br /&gt;
File:Screenshot 2023-03-24 122521.png | link=http://borreliabase.org/~wgqiu/ospC-sequencing | OspC amplicon sequencing from ticks&lt;br /&gt;
File:Screenshot 2023-03-24 122340.png | link=http://borreliabase.org/~wgqiu/asm-challenge | Genomic epidemiology of a Salmonella outbreak (ASM Challenge)&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Curricular Development &amp;amp; Bioinformatics/QuBi Advising ==&lt;br /&gt;
*QuBi advisors: Weigang Qiu, Ntino Krampis, Rabindra Mandal (Biology); Saad Mneimeih, Lei Xie (CS); Akira Kawamura (Chem); Dana Sylvan (Math &amp;amp; Stats) &lt;br /&gt;
**Permission for non-Biology majors to take BIOL203 &amp;amp; BIOL425, every Spring&lt;br /&gt;
**Collect names,  major, and IDs to send to course coordinator to grant permission. Waive BIOL10200 pre-reqs for taking BIOL203.&lt;br /&gt;
*Curricular resources:&lt;br /&gt;
**[http://biology.hunter.cuny.edu/index.php?option=com_content&amp;amp;view=article&amp;amp;id=66&amp;amp;Itemid=73 Biology courses and pre-reqs]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/BIO1-BA Hunter Biology Major 1 (including the Bioinformatics Option)]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/CHEM2-BA Hunter Chemistry Major 2 (including the Bioinformatics Option)]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/COMPSCI-BA Hunter Computer Science (including Bioinformatics Concentration)]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/MATH-BA Hunter Mathematics (including the Quantitative Biology Concentration)]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/STATS-BA Hunter Statistics (including the Quantitative Biology Concentration)]&lt;br /&gt;
*QuBi advising:&lt;br /&gt;
**Declaration of Bioinformatics concentration: In-person advising to work out the semester-by-semester courses&lt;br /&gt;
**Approve on department spreadsheet (or send email to  &amp;quot;Samantha Sheppard-Lahiji&amp;quot; and &amp;quot;HTR Bio&amp;quot; &amp;lt;biology@hunter.cuny.edu&amp;gt;)&lt;br /&gt;
**Students should take Bioinformatics-specific electives (8 cred; see Hunter Catalog below), &#039;&#039;&#039;not general electives&#039;&#039;&#039;&lt;br /&gt;
***Examples: Anthrop302 (3 cr); Chem333 (3 cr); BIOL47119 &amp;amp; BIOL47120 (3cr); BIOL48002 (2 cr)&lt;br /&gt;
**Students need to take BIOL48002 (2 cr), which counts towards as research credit, to graduate as honors&lt;br /&gt;
*General advising:&lt;br /&gt;
**~40 students every semester. Send out emails to students. Go through student courses by Email or by appointment&lt;br /&gt;
**Recommend new math courses: &#039;&#039;&#039;MATH15200 &amp;amp; STAT21350&#039;&#039;&#039;&lt;br /&gt;
*Hosting QuBi students in lab&lt;br /&gt;
**This is to enhance the informatics and coding skills of our students&lt;br /&gt;
**Students should register and get &#039;&#039;&#039;BIOL48002&#039;&#039;&#039; credits, which counts towards their elective credits &amp;amp; eligibility for honors&lt;br /&gt;
**3-5 students per semester&lt;br /&gt;
*Outside research opportunities&lt;br /&gt;
**MIT Quantitative Workshop (first week of January, in Boston). Coordination with CS (Saad &amp;amp; Susan Epstein) in Fall&lt;br /&gt;
**Simons Foundation/Flatiron Institute Center for Computational Biology (CCB) Internship program. Open House in Spring&lt;br /&gt;
&lt;br /&gt;
==Course/Lecture syllabi==&lt;br /&gt;
*[[NYRaMP-Informatics-2026|NYRaMP Workshop (August 2026, by Brandon Ely)]]&lt;br /&gt;
*[[NYRaMP-Informatics-2025|NYRaMP Workshop (August 2025, by Brandon Ely)]]&lt;br /&gt;
*[[Computational Genomics Summer 2026 | KIZ Genomics Workshop, Summer 2026]]&lt;br /&gt;
*[[Computational Genomics (KIZ, Fall 2024)]]&lt;br /&gt;
*[[NYRaMP-Informatics-2024|NYRaMP Workshop (August 2024)]]&lt;br /&gt;
*BIOL47120 BioMedical Genomics (Spring 2024). Tutorials: [https://borreliabase.org/~wgqiu/tutorial-markdown.html R Markdown] [https://borreliabase.org/~wgqiu/cluster-analysis.html Cluster analysis] [https://borreliabase.org/~wgqiu/scRNA-analysis.html single-cell RNA-seq]&lt;br /&gt;
*BIOL425 Computational Molecular Biology (Spring, 2023). [https://github.com/weigangq/CSB-BIOL425/tree/master/lecture-materials Lecture material on github]&lt;br /&gt;
*BIOL714 Cell Biology: [http://borreliabase.org/~wgqiu/r-demo-2024.html R Demo (Spring 2024)] [http://borreliabase.org/~wgqiu/r-demo-2023.html R Demo (Spring 2023)]&lt;br /&gt;
*QuBi module: [[QuBi/module/bio203-lab12—2022|BIOL20300 Molecular Genetics, Lab 12 (2023)]]&lt;br /&gt;
*QuBi module: [[QuBi/modules/biol203-geno-pheno-association-2022|BIOL20300 Molecular Genetics, Lab 13 (2022)]]&lt;br /&gt;
*QuBi module: [[QuBi/modules/biol303|BIOL30300 Cell Biology, Bioinformatics Lab (transcriptome analysis)]]&lt;br /&gt;
*[[BigData 2020|Big Data (Summer, 2020)]]&lt;br /&gt;
*[[BioMed-R-2020|BIOL47120 Biomedical Genomics II (Spring, 2020)]] &lt;br /&gt;
**[http://borreliabase.org/~wgqiu/tutorial-markdown.html Tutorial: R Markdown (Spring 2024)]&lt;br /&gt;
**[http://borreliabase.org/~wgqiu/cluster-analysis.html Tutorial: Cluster analysis (Spring 2024)]&lt;br /&gt;
**[http://borreliabase.org/~wgqiu/scRNA-analysis.html Tutorial: single-cel transcriptome analysis (Spring 2024)]&lt;br /&gt;
*[[Biol425 2020|BIOL425 Computational Molecular Biology (Spring, 2020)]]&lt;br /&gt;
*[[Biol375 2019|BIOL37500, Molecular Evolution (Fall, 2019)]]&lt;br /&gt;
*[[Southwest-University|Southwest University R course (Summer, 2019)]]&lt;br /&gt;
*[[Biol20N02 2017|Analysis of Biological Data (Spring, 2017)]]&lt;br /&gt;
*[[Bioinformatics_Workshop_2014|Bioinformatics Workshop (Summer, 2014)]]&lt;br /&gt;
&lt;br /&gt;
==SARS-CoV-2 genome evolution==&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;3&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File:Cov-fig1.jpg | Akther, Bezrucenkovas, Sulkow, Panlasigui, Qiu, Di (April, 2020). &amp;quot;CoV Genome Tracker: tracing genomic footprints of Covid-19 pandemic&amp;quot;. &#039;&#039;&#039;&#039;&#039;[https://www.biorxiv.org/content/biorxiv/early/2020/04/14/2020.04.10.036343.full.pdf BioRxiv]&#039;&#039;&#039;&#039;&#039;; [https://cov.genometracker.org/ Web app: SARS-CoV-2 Genome Tracker]; Github: https://github.com/weigangq/cov-browser&lt;br /&gt;
File:Rec-fig2.png | Akther, Li, Martin, Di, Sulkow, Pante, Bezrucenlovas, Luft, Qiu (May, 2020). &amp;quot;Origin, recombination, and missed opprotunities:  a genomic perspective of the first 100 days of COVID-19 pandemic&amp;quot;. (Unpublished).&lt;br /&gt;
File:Cov-fig3-trace.png | Akther, Bezrucenlovas, Li, Sulkow, Di, Pante, Martin, Luft, Qiu (Sep, 2021). &amp;quot;Following the Trail of One Million Genomes: Footprints of SARS-CoV-2 Adaptation to Humans&amp;quot;. &#039;&#039;&#039;&#039;&#039;[https://www.biorxiv.org/content/biorxiv/early/2021/05/10/2021.05.07.443114.full.pdf BioRxiv link]&#039;&#039;&#039;&#039;&#039;: . Github:  https://github.com/weigangq/cov-db&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Lab Resources &amp;amp; Protocols==&lt;br /&gt;
*OneRF usage (Unofficial; login needed): https://sites.google.com/view/onerf/home&lt;br /&gt;
*[https://runestone.academy/ns/books/published/thinkcspy/index.html How to think like a computer scientist: An interactive Python programming book]&lt;br /&gt;
* Nanopore sequencing protocols&lt;br /&gt;
**DNA barcoding: https://nanopore4edu.org/latest/annotated_experiments/dna_barcoding/&lt;br /&gt;
**Yeast genomes: https://nanoporetech.com/document/extraction-method/yeast-dna&lt;br /&gt;
*[[Monte Carlo Club]]&lt;br /&gt;
*[[NY-RaMP Mentoring]]&lt;br /&gt;
*Borreliella genome sequencing consortium: Weekly meetings (Tu @11): Since Jan 2023&lt;br /&gt;
*Borreliella diagnostic antigens (Fall 2023-Fall 2027):&lt;br /&gt;
**Zoom call (Jan 23, 2024)&lt;br /&gt;
** Next meeting: March 23, 2024&lt;br /&gt;
*Qiu lab network [[First Time Guide|first-time user guide]]&lt;br /&gt;
*Qiu lab Github repositories: https://github.com/weigangq/?tab=repositories&lt;br /&gt;
*[[Mini-Tutorals|Mini-Protocols]] (frequently used computer codes and pipelines)&lt;br /&gt;
*Python tutorial: https://wiki.genometracker.org/~weigang/Intro_to_Python.html&lt;br /&gt;
* [[Tick protocol|ick handling protocols]]&lt;br /&gt;
*[[A Primer on the Cluster System at Hunter|Hunter HPC Usage]]&lt;br /&gt;
*[https://r4ds.hadley.nz/ R for Data Science (2e)], (2024) by Wickham, Grolemund &amp;amp; Çetinkaya-Rundel ([https://bookdown.org/ Bookdown version])&lt;br /&gt;
*Borrelia Genome Consortium:  https://www.ncbi.nlm.nih.gov/bioproject/PRJNA431102/&lt;br /&gt;
*Canadian Bbsl genome assemblies: https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA1130942&lt;br /&gt;
* Nanopore sequencing resources:&lt;br /&gt;
**eBook: https://store.nanoporetech.com/us/minion.html&lt;br /&gt;
&lt;br /&gt;
==Wiki Help==&lt;br /&gt;
*[https://www.mediawiki.org/wiki/Special:MyLanguage/Manual:Configuration_settings Configuration settings list]&lt;br /&gt;
*[https://www.mediawiki.org/wiki/Special:MyLanguage/Manual:FAQ MediaWiki FAQ]&lt;br /&gt;
*[https://lists.wikimedia.org/postorius/lists/mediawiki-announce.lists.wikimedia.org/ MediaWiki release mailing list]&lt;br /&gt;
*[https://www.mediawiki.org/wiki/Special:MyLanguage/Localisation#Translation_resources Localise MediaWiki for your language]&lt;br /&gt;
*[https://www.mediawiki.org/wiki/Special:MyLanguage/Manual:Combating_spam Learn how to combat spam on your wiki]&lt;br /&gt;
*Consult the [[mediawikiwiki:Special:MyLanguage/Help:Contents|User&#039;s Guide]] for information on using the wiki software.&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=NYRaMP-Informatics-2026&amp;diff=6704</id>
		<title>NYRaMP-Informatics-2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=NYRaMP-Informatics-2026&amp;diff=6704"/>
		<updated>2026-07-02T10:47:24Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: Created page with &amp;quot;&amp;lt;center&amp;gt;&amp;#039;&amp;#039;&amp;#039;NYRaMP Informatics Workshop&amp;#039;&amp;#039;&amp;#039;&amp;lt;/center&amp;gt; &amp;lt;center&amp;gt;August 2026, Tuesdays 9-12, DNA Learning Center&amp;lt;/center&amp;gt; &amp;lt;center&amp;gt;&amp;#039;&amp;#039;&amp;#039;Instructors:&amp;#039;&amp;#039;&amp;#039; Dr. Brandon Ely&amp;lt;/center&amp;gt; &amp;lt;center&amp;gt; {| class=&amp;quot;wikitable&amp;quot; |- ! MA plot !! Volcano plot !! Heat map |- |  fold change (y-axis) vs. total expression levels (x-axis) ||   p-value (y-axis) vs. fold change (x-axis) ||   File:GeneExp3.jpeg|300px|thumbnail| gene...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;&#039;&#039;&#039;NYRaMP Informatics Workshop&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;August 2026, Tuesdays 9-12, DNA Learning Center&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Instructors:&#039;&#039;&#039; Dr. Brandon Ely&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! MA plot !! Volcano plot !! Heat map&lt;br /&gt;
|-&lt;br /&gt;
| [[File:GeneExp1.jpeg|300px|thumbnail| fold change (y-axis) vs. total expression levels (x-axis)]] || &lt;br /&gt;
[[File:GeneExp2.jpeg|300px|thumbnail| p-value (y-axis) vs. fold change (x-axis)]]&lt;br /&gt;
|| &lt;br /&gt;
 [[File:GeneExp3.jpeg|300px|thumbnail| genes significantly down or up-regulated (at p&amp;lt;1e-4)]]&lt;br /&gt;
|}&lt;br /&gt;
&amp;lt;/center&amp;gt;&lt;br /&gt;
==Overview==&lt;br /&gt;
A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and rapid DNA and RNA-sequencing technologies, biomedical sciences are undergoing a rapid &amp;amp; irreversible transformation into a highly data-intensive field, that requires familiarity with concepts in both biological, computational, and statistical sciences.  &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life Sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as statistics. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this workshop students will be able to: &lt;br /&gt;
* Manipulate &amp;amp; visualize biological data sets with R &amp;amp; Rstudio&lt;br /&gt;
* Simulate the Central Dogma&lt;br /&gt;
* Analyze transcriptome data&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Cloud R account (free): https://posit.cloud/; Join the shared work space &amp;quot;NYRaMP-Informatics&amp;quot;&lt;br /&gt;
* For your own computer, download the desktop version: https://posit.co/download/rstudio-desktop/&lt;br /&gt;
* Textbook: [http://r4all.org/#about Introduction to R for Biologists]&lt;br /&gt;
* Download: [http://www.r4all.org/books/datasets R datasets]&lt;br /&gt;
* A reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm &amp;amp; Grolemund)]&lt;br /&gt;
&lt;br /&gt;
==Week 1. Aug 4==&lt;br /&gt;
* Pre-test: visualization, interpretation, and stats. Download file: [[File:Pre-test.pdf|thumb]]&lt;br /&gt;
* Computer/Cloud setup &amp;amp; software download/installation&lt;br /&gt;
* R Tutorial 1. Getting started: Basics: interface, packages, variables, objects, functions. Download slides: [[File:NYRamP_bioinformatics_1_slides.pdf|thumb]]&lt;br /&gt;
* Session 1 R code: Basic R syntax, working with vectors, and using functions&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&#039;bash&#039;&amp;gt;&lt;br /&gt;
##### Practice 1 - Together #####&lt;br /&gt;
&lt;br /&gt;
# TASK 1: define a variable that is your name &lt;br /&gt;
MyName &amp;lt;- &#039;Brandon&#039;&lt;br /&gt;
print(MyName)&lt;br /&gt;
paste(&#039;My name is&#039;,MyName, sep = &#039; &#039;)&lt;br /&gt;
&lt;br /&gt;
# TASK 2: output the 3rd and 4th letters in your name using substr function&lt;br /&gt;
substr(MyName, start = 3, stop = 4)&lt;br /&gt;
&lt;br /&gt;
# TASK 3: create a vector of the names of all of your Ramp cohort members &lt;br /&gt;
roster &amp;lt;- c(&#039;Amalya&#039;, &#039;Danny&#039;, &#039;Lorelei&#039;, &#039;Dylan&#039;, &#039;Hadley&#039;, &#039;Brynn&#039;, &#039;Elliot&#039;, &#039;Theo&#039;)&lt;br /&gt;
&lt;br /&gt;
# Task 4: check if any of the names have the letters &amp;quot;ic&amp;quot; in them&lt;br /&gt;
grepl(&#039;an&#039;, roster, ignore.case = F)&lt;br /&gt;
&lt;br /&gt;
# Task 5: randomly select 3 names from the roster&lt;br /&gt;
sample(roster, size = 3, replace = FALSE)&lt;br /&gt;
&lt;br /&gt;
# TASK 6: combine tasks 4-5&lt;br /&gt;
grepl(&#039;an&#039;, sample(roster, size = 3, replace = F), ignore.case = F)&lt;br /&gt;
&lt;br /&gt;
##### Practice 2 - Independent  #####&lt;br /&gt;
&lt;br /&gt;
library(stringr)&lt;br /&gt;
&lt;br /&gt;
# TASK 1: create a character vector for the DNA nucleotides &lt;br /&gt;
Nucleotides &amp;lt;- c(&#039;A&#039;, &#039;T&#039;, &#039;C&#039;, &#039;G&#039;)&lt;br /&gt;
&lt;br /&gt;
# TASK 2: use the &amp;quot;sample&amp;quot; function on your nuc vector to create a DNA sequence of length 200&lt;br /&gt;
DNAseq &amp;lt;- sample(Nucleotides, 200, replace = TRUE)&lt;br /&gt;
&lt;br /&gt;
# collapse the vector into a single string &lt;br /&gt;
DNAseq2 &amp;lt;- paste(DNAseq, collapse = &#039;&#039;)&lt;br /&gt;
&lt;br /&gt;
# TASK 3: Find out if your sequence contains start codons using the &amp;quot;grepl&amp;quot; function; if output = FALSE, start over&lt;br /&gt;
grepl(&#039;ATG&#039;, DNAseq2, ignore.case = F)&lt;br /&gt;
&lt;br /&gt;
# TASK 4: Find all locations of start codons within your sequence using &amp;quot;str_locate_all&amp;quot; function&lt;br /&gt;
str_locate_all(string = DNAseq2, pattern = &#039;ATG&#039;)&lt;br /&gt;
&lt;br /&gt;
# TASK 5: use &amp;quot;substring&amp;quot; function to confirm coordinates are actually &amp;quot;ATG&amp;quot;&lt;br /&gt;
&#039;ATG&#039; == substr(DNAseq2, start = 96, stop = 98)&lt;br /&gt;
&lt;br /&gt;
# TASK 6: calculate nucleotide % composition using &amp;quot;str_count&amp;quot; function&lt;br /&gt;
str_count(string = DNAseq2, pattern = &#039;A&#039;) / 200 * 100&lt;br /&gt;
&lt;br /&gt;
### alternate more advanced way for tasks 5 and 6 ###&lt;br /&gt;
&lt;br /&gt;
coords &amp;lt;- str_locate_all(string = DNAseq2, pattern = &#039;ATG&#039;)&lt;br /&gt;
&lt;br /&gt;
for (i in 1:nrow(coords[[1]])) {&lt;br /&gt;
  start &amp;lt;- coords[[1]][i, 1]&lt;br /&gt;
  print(&#039;ATG&#039; == substr(DNAseq2, start = start, stop = start+2))&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
for (nuc in Nucleotides) {&lt;br /&gt;
  print(paste(nuc,&#039; = &#039;, str_count(string = DNAseq2, pattern = nuc)/nchar(DNAseq2)*100, sep = &#039;&#039;))&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Week 2. Aug 11==&lt;br /&gt;
* R Tutorial 2. Data manipulation. Download slides: [[File:NYRamP_bioinformatics_2_slides.pdf|thumb]]&lt;br /&gt;
* Session 2 R code:&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&#039;bash&#039;&amp;gt;&lt;br /&gt;
library(tidyverse)&lt;br /&gt;
library(broom)&lt;br /&gt;
&lt;br /&gt;
##### loading a data table #####&lt;br /&gt;
df &amp;lt;- read.table(&#039;http://wiki.genometracker.org/~weigang/datasets-master/compensation.csv&#039;, sep = &#039;,&#039;, header = T)&lt;br /&gt;
# this loads the data&lt;br /&gt;
&lt;br /&gt;
##### quick look at data #####&lt;br /&gt;
dim(df)&lt;br /&gt;
nrow(df)&lt;br /&gt;
ncol(df)&lt;br /&gt;
names(df)&lt;br /&gt;
head(df)&lt;br /&gt;
glimpse(df)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
##### use slice to grab specific rows #####&lt;br /&gt;
slice(df, 11:15)&lt;br /&gt;
slice_head(df, n=5)&lt;br /&gt;
slice_tail(df, n=8)&lt;br /&gt;
slice_sample(df, n=20)&lt;br /&gt;
slice_max(df, Root, n=3)&lt;br /&gt;
slice_min(df, Root, n=3)&lt;br /&gt;
slice_min(df, Grazing)&lt;br /&gt;
&lt;br /&gt;
##### arrange function #####&lt;br /&gt;
arrange(df, Fruit)&lt;br /&gt;
arrange(df, desc(Fruit))&lt;br /&gt;
arrange(df, Grazing)&lt;br /&gt;
arrange(df, desc(Grazing))&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
##### filter #####&lt;br /&gt;
filter(df, Grazing == &#039;Grazed&#039;)&lt;br /&gt;
filter(df, Root &amp;lt; 6.5)&lt;br /&gt;
filter(df, Grazing == &#039;Ungrazed&#039; &amp;amp; Fruit &amp;gt; 50)&lt;br /&gt;
filter(df, Fruit &amp;gt; 50 | Root &amp;gt;8)&lt;br /&gt;
filter(df,  Grazing == &#039;Ungrazed&#039; &amp;amp; ( Fruit &amp;gt; 50 | Root &amp;gt;8))&lt;br /&gt;
&lt;br /&gt;
##### piping #####&lt;br /&gt;
df %&amp;gt;% #start by calling my dataframe&lt;br /&gt;
  select(-Root) %&amp;gt;% # removing the Root data&lt;br /&gt;
  filter(Fruit &amp;gt; 50) %&amp;gt;% # keep only fruit with value greater than 50&lt;br /&gt;
  arrange(Grazing, Fruit) # arrange first by grazed vs ungrazed, then by fruit size least to greatest&lt;br /&gt;
&lt;br /&gt;
##### computations #####&lt;br /&gt;
&lt;br /&gt;
# operations on an entire column&lt;br /&gt;
mean(df$Fruit)&lt;br /&gt;
median(df$Root)&lt;br /&gt;
sum(df$Fruit)&lt;br /&gt;
&lt;br /&gt;
# operations on individual values in the column &lt;br /&gt;
log2(df$Fruit)&lt;br /&gt;
df$Fruit / 2&lt;br /&gt;
&lt;br /&gt;
# make a new column by defining and assigning&lt;br /&gt;
new_df &amp;lt;- df&lt;br /&gt;
new_df$test &amp;lt;- df$Root + df$Fruit&lt;br /&gt;
&lt;br /&gt;
# making a new column based on a computation of data in the df using mutate&lt;br /&gt;
new_df &amp;lt;- df %&amp;gt;%&lt;br /&gt;
  mutate(log_fruit = log2(Fruit)) %&amp;gt;%&lt;br /&gt;
  mutate(combined_mass = Fruit + Root)&lt;br /&gt;
&lt;br /&gt;
##### summary statistics #####&lt;br /&gt;
summary(df)&lt;br /&gt;
&lt;br /&gt;
# pipe to get summary by specific variables&lt;br /&gt;
df %&amp;gt;% &lt;br /&gt;
  filter(Grazing == &#039;Grazed&#039;) %&amp;gt;%&lt;br /&gt;
  summary()&lt;br /&gt;
&lt;br /&gt;
# get the stats you want, with grouping &lt;br /&gt;
df %&amp;gt;% &lt;br /&gt;
  group_by(Grazing) %&amp;gt;%&lt;br /&gt;
  summarise(&lt;br /&gt;
    mean_fruit = mean(Fruit),&lt;br /&gt;
    median_fruit = median(Fruit),&lt;br /&gt;
    sd_fruit = sd(Fruit),&lt;br /&gt;
    mean_root = mean(Root),&lt;br /&gt;
    median_root = median(Root),&lt;br /&gt;
    sd_root = sd(Root)&lt;br /&gt;
  )&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
##### rotating a dataframe #####&lt;br /&gt;
# add an id so values are in unique rows&lt;br /&gt;
df$id &amp;lt;- as.character(1:length(df$Grazing)) &lt;br /&gt;
&lt;br /&gt;
# convert to long format&lt;br /&gt;
df_long &amp;lt;- df %&amp;gt;% pivot_longer(cols = c(Fruit, Root), values_to = &#039;value&#039;, names_to = &#039;type&#039;)&lt;br /&gt;
&lt;br /&gt;
# you can still group and summarize like you can with the wide format&lt;br /&gt;
df_long %&amp;gt;% &lt;br /&gt;
  group_by(type, Grazing) %&amp;gt;%&lt;br /&gt;
  summarise(&lt;br /&gt;
    mean = mean(value)&lt;br /&gt;
  )&lt;br /&gt;
&lt;br /&gt;
# convert a long format df to wide&lt;br /&gt;
df_wide &amp;lt;- df_long %&amp;gt;% pivot_wider(names_from = type, values_from = value)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
##### IRIS DATASET PRACTICE #####&lt;br /&gt;
data(&amp;quot;iris&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
# 1 - Find the dimensions of the dataset&lt;br /&gt;
dim(iris)&lt;br /&gt;
&lt;br /&gt;
paste(&#039;The iris dataset has&#039;,dim(iris)[1],&#039;rows and&#039;,dim(iris)[2],&#039;columns.&#039;, sep = &#039; &#039;)&lt;br /&gt;
&lt;br /&gt;
# 2 - List the variables and their data types&lt;br /&gt;
glimpse(iris)&lt;br /&gt;
&lt;br /&gt;
# 3 - Summarize the variables&lt;br /&gt;
summary(iris)&lt;br /&gt;
&lt;br /&gt;
# 4 - Get the last 10 observations of the dataset&lt;br /&gt;
tail(iris, n = 10)&lt;br /&gt;
slice_tail(iris, n = 10)&lt;br /&gt;
&lt;br /&gt;
# 5 - Select only the first four columns (remove the “species” column)&lt;br /&gt;
select(iris, 1:4)&lt;br /&gt;
iris %&amp;gt;% select(1:4)&lt;br /&gt;
select(iris, -Species)&lt;br /&gt;
&lt;br /&gt;
# 6 - Filter rows by species, retain only rows from one species (e.g., “setosa”)&lt;br /&gt;
filter(iris, Species == &#039;setosa&#039;)&lt;br /&gt;
iris %&amp;gt;% filter(Species == &#039;setosa&#039;)&lt;br /&gt;
&lt;br /&gt;
# 7 - Filter rows by a cutoff value (e.g., “Sepal.Length &amp;gt;= 4”)&lt;br /&gt;
iris %&amp;gt;% filter(Sepal.Length &amp;gt;= 5)&lt;br /&gt;
&lt;br /&gt;
# 8 - Add a column by taking the log10 of “Sepal.Lengh”&lt;br /&gt;
iris %&amp;gt;% mutate(log10_sl = log10(Sepal.Length))&lt;br /&gt;
&lt;br /&gt;
# 9 - What are the medians of the variable “Sepal.Length” for each species?&lt;br /&gt;
iris %&amp;gt;%&lt;br /&gt;
  group_by(Species) %&amp;gt;%&lt;br /&gt;
  summarise(&lt;br /&gt;
    median_sl = median(Sepal.Length)&lt;br /&gt;
  )&lt;br /&gt;
&lt;br /&gt;
# 10 - Count how many samples for each species&lt;br /&gt;
count(iris, Species)&lt;br /&gt;
iris %&amp;gt;% count(Species)&lt;br /&gt;
&lt;br /&gt;
# 11 - Convert the table to long format&lt;br /&gt;
iris %&amp;gt;%&lt;br /&gt;
  pivot_longer(cols = c(Sepal.Length,Sepal.Width,Petal.Length,Petal.Width), &lt;br /&gt;
               names_to = &#039;structure&#039;,&lt;br /&gt;
               values_to = &#039;measurement&#039;)&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Week 3. Aug 18==&lt;br /&gt;
* R Tutorial 3. Data visualization. Lecture slides: [[File:NYRamP_bioinformatics_3_slides.pdf|thumb]]&lt;br /&gt;
* Session 3 R code:&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&#039;bash&#039;&amp;gt;&lt;br /&gt;
library(tidyverse)&lt;br /&gt;
library(broom)&lt;br /&gt;
##### compensation dataset viz #####&lt;br /&gt;
&lt;br /&gt;
df &amp;lt;- read.table(&#039;http://wiki.genometracker.org/~weigang/datasets-master/compensation.csv&#039;, sep = &#039;,&#039;, header = T)&lt;br /&gt;
&lt;br /&gt;
# boxplot for root or fruit measurements&lt;br /&gt;
df %&amp;gt;%&lt;br /&gt;
  ggplot(aes(x = Grazing, y = Root, fill = Grazing)) +&lt;br /&gt;
  geom_boxplot() +&lt;br /&gt;
  #geom_point() &lt;br /&gt;
  geom_jitter(color = &#039;blue&#039;)&lt;br /&gt;
&lt;br /&gt;
# scatter plot with regression line&lt;br /&gt;
df %&amp;gt;%&lt;br /&gt;
  ggplot(aes(x = Root, y = Fruit, color = Grazing)) +&lt;br /&gt;
  geom_point() +&lt;br /&gt;
  geom_smooth(method = &#039;lm&#039;)&lt;br /&gt;
&lt;br /&gt;
# bar plot for root or fruit measurement &lt;br /&gt;
df %&amp;gt;% &lt;br /&gt;
  ggplot(aes(x = Grazing, y = Fruit, fill = Grazing)) +&lt;br /&gt;
  geom_bar(stat = &#039;identity&#039;)&lt;br /&gt;
&lt;br /&gt;
# we can do more with a long data format&lt;br /&gt;
df_long &amp;lt;- df %&amp;gt;% pivot_longer(cols = c(Fruit, Root), values_to = &#039;value&#039;, names_to = &#039;type&#039;)&lt;br /&gt;
&lt;br /&gt;
# box plot showing root and fruit together&lt;br /&gt;
df_long %&amp;gt;%&lt;br /&gt;
  ggplot(aes(x = type, y = value, fill = Grazing)) +&lt;br /&gt;
  geom_boxplot() +&lt;br /&gt;
  geom_jitter() +&lt;br /&gt;
  facet_wrap(~type, scales = &#039;free&#039;)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
data(iris)&lt;br /&gt;
##### compensation dataset hypothesis testing #####&lt;br /&gt;
# is there a correlation between root and fruit variables?&lt;br /&gt;
cor(df$Root, df$Fruit)&lt;br /&gt;
&lt;br /&gt;
# lets group by to get a better idea&lt;br /&gt;
df %&amp;gt;% &lt;br /&gt;
  group_by(Grazing) %&amp;gt;%&lt;br /&gt;
  summarise(pearson = cor(Root, Fruit))&lt;br /&gt;
&lt;br /&gt;
# is the mean Fruit size the same for grazed and ungrazed? &lt;br /&gt;
# t test&lt;br /&gt;
t.test(df$Fruit ~ df$Grazing)&lt;br /&gt;
&lt;br /&gt;
# tidy for results in a neat dataframe&lt;br /&gt;
res &amp;lt;- tidy(t.test(df$Fruit ~ df$Grazing))&lt;br /&gt;
&lt;br /&gt;
# another way of doing t.test + tidy&lt;br /&gt;
t_test &amp;lt;- df %&amp;gt;%&lt;br /&gt;
  t.test(Fruit ~ Grazing, data = .) %&amp;gt;%&lt;br /&gt;
  tidy()&lt;br /&gt;
&lt;br /&gt;
# linear model - can these variables predict each other?&lt;br /&gt;
# can root size predict fruit mass? &lt;br /&gt;
lm_simple &amp;lt;- lm(Fruit ~ Root, data = df)&lt;br /&gt;
lm_interact &amp;lt;- lm(Fruit ~ Root * Grazing, data = df)&lt;br /&gt;
&lt;br /&gt;
tidy(lm_simple)&lt;br /&gt;
tidy(lm_interact)&lt;br /&gt;
&lt;br /&gt;
summary(lm(Fruit ~ Root, data = df))&lt;br /&gt;
summary(lm(Fruit ~ Root * Grazing, data = df))&lt;br /&gt;
##### iris dataset Viz #####&lt;br /&gt;
&lt;br /&gt;
# simple 1 variable boxplot&lt;br /&gt;
iris %&amp;gt;%&lt;br /&gt;
  ggplot(aes(x = Species, y = Sepal.Length, fill = Species)) +&lt;br /&gt;
  #geom_boxplot() +&lt;br /&gt;
  geom_violin() +&lt;br /&gt;
  #geom_point()&lt;br /&gt;
  geom_jitter()  &lt;br /&gt;
#geom_bar(stat = &#039;identity&#039;) &lt;br /&gt;
&lt;br /&gt;
# scatter plot for 2 numerical variables&lt;br /&gt;
iris %&amp;gt;%&lt;br /&gt;
  ggplot(aes(x = Sepal.Length, y = Petal.Length, color = Species)) +&lt;br /&gt;
  geom_point() +&lt;br /&gt;
  geom_smooth(method = &amp;quot;lm&amp;quot;, se = T)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
# lets get more information at once!&lt;br /&gt;
# convert to long format&lt;br /&gt;
df_long &amp;lt;- iris %&amp;gt;%&lt;br /&gt;
  pivot_longer(cols = c(Sepal.Length,Sepal.Width,Petal.Length,Petal.Width), &lt;br /&gt;
               names_to = &#039;structure&#039;,&lt;br /&gt;
               values_to = &#039;measurement&#039;)&lt;br /&gt;
&lt;br /&gt;
df_long %&amp;gt;%&lt;br /&gt;
  #filter(structure == &#039;Sepal.Length&#039;) %&amp;gt;%&lt;br /&gt;
  ggplot(aes(x = Species, y = measurement, fill = Species)) +&lt;br /&gt;
  geom_boxplot() +&lt;br /&gt;
  geom_jitter(alpha = 0.5) +&lt;br /&gt;
  facet_wrap(~structure, scales = &#039;free_y&#039;)&lt;br /&gt;
&lt;br /&gt;
df_long %&amp;gt;%&lt;br /&gt;
  ggplot(aes(x = Species, y = measurement, fill = Species)) +&lt;br /&gt;
  geom_bar(stat = &#039;identity&#039;) +&lt;br /&gt;
  facet_wrap(~structure)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
##### iris dataset hypothesis testing #####&lt;br /&gt;
# 1 - is there a correlation between sepal length and sepal width?&lt;br /&gt;
cor(iris$Sepal.Length, iris$Sepal.Width)&lt;br /&gt;
cor(iris$Sepal.Length, iris$Petal.Length)&lt;br /&gt;
&lt;br /&gt;
# 2 - is there a difference in mean sepal length among the different species?&lt;br /&gt;
# anova to compare all three&lt;br /&gt;
anova_res &amp;lt;- aov(Sepal.Length ~ Species, data = iris)&lt;br /&gt;
# look at results&lt;br /&gt;
summary(anova_res)&lt;br /&gt;
# post hoc test to see which groups differ&lt;br /&gt;
TukeyHSD(anova_res)&lt;br /&gt;
&lt;br /&gt;
# we can do the similar things with t-tests &lt;br /&gt;
df &amp;lt;- iris %&amp;gt;% &lt;br /&gt;
  filter(Species == &#039;setosa&#039; | Species == &#039;virginica&#039;)&lt;br /&gt;
&lt;br /&gt;
t.test(Sepal.Length ~ Species, df)$p.value&lt;br /&gt;
&lt;br /&gt;
t_res &amp;lt;- t.test(Sepal.Length ~ Species, data = df)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Demo: Microbiome data analysis==&lt;br /&gt;
* Mock gene expression data set: [[File:sample_gene_exp_dataset.csv|thumb]]&lt;br /&gt;
* DNA barcoding dataset, NYC soil microbiome [[File:abundance_table_species.tsv|thumb]]&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&#039;bash&#039;&amp;gt;&lt;br /&gt;
library(tidyverse)&lt;br /&gt;
library(pheatmap)&lt;br /&gt;
library(vegan)&lt;br /&gt;
&lt;br /&gt;
setwd(&#039;/Users/brandonely/Desktop/informatics_bootcamp/summer_25/&#039;)&lt;br /&gt;
&lt;br /&gt;
##### READ DATA #####&lt;br /&gt;
df &amp;lt;- read.table(&#039;abundance_table_species.tsv&#039;, sep = &#039;\t&#039;, header = T)&lt;br /&gt;
&lt;br /&gt;
# take a quick look and what we&#039;re working with&lt;br /&gt;
glimpse(df)&lt;br /&gt;
head(df)&lt;br /&gt;
&lt;br /&gt;
##### CLEAN UP DF #####&lt;br /&gt;
&lt;br /&gt;
# lets tackle that tax column and separate it into separate classification levels&lt;br /&gt;
df2 &amp;lt;- df %&amp;gt;% &lt;br /&gt;
  separate(tax, &lt;br /&gt;
           into = c(&#039;domain&#039;, &lt;br /&gt;
                    &#039;kingdom&#039;, &lt;br /&gt;
                    &#039;phylum&#039;, &lt;br /&gt;
                    &#039;class&#039;, &lt;br /&gt;
                    &#039;order&#039;, &lt;br /&gt;
                    &#039;family&#039;, &lt;br /&gt;
                    &#039;genus&#039;, &lt;br /&gt;
                    &#039;species&#039;), &lt;br /&gt;
            sep = &#039;;&#039;, &lt;br /&gt;
            fill = &#039;warn&#039;)&lt;br /&gt;
&lt;br /&gt;
# clean up the genus / species column &lt;br /&gt;
df2 &amp;lt;- df2 %&amp;gt;%&lt;br /&gt;
  separate(species,into =c(&#039;genus2&#039;, &#039;species&#039;),sep = &#039; &#039;,fill = &#039;warn&#039;) %&amp;gt;%&lt;br /&gt;
  select(c(phylum, genus, species, 9:25)) &lt;br /&gt;
&lt;br /&gt;
# remove last row of df to avoid downstream issues &lt;br /&gt;
df2 &amp;lt;- df2[-nrow(df2), ]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
##### PHYLUM LEVEL COMMUNITY COMPOSITION BAR PLOTS #####&lt;br /&gt;
&lt;br /&gt;
# for each phylum, get total read counts per sample&lt;br /&gt;
df_phylum &amp;lt;- df2 %&amp;gt;%&lt;br /&gt;
  select(-c(genus, species)) %&amp;gt;%  # remove uneeded columns&lt;br /&gt;
  group_by(phylum) %&amp;gt;%&lt;br /&gt;
  summarise(across(starts_with(&amp;quot;barcode&amp;quot;), sum), .groups=&amp;quot;drop&amp;quot;) # sum counts for each phylum&lt;br /&gt;
&lt;br /&gt;
# convert df to long format &lt;br /&gt;
df_phylum &amp;lt;- df_phylum %&amp;gt;% &lt;br /&gt;
  pivot_longer(-phylum, names_to=&amp;quot;sample&amp;quot;, values_to=&amp;quot;abundance&amp;quot;) # *why the -phylum here? &lt;br /&gt;
&lt;br /&gt;
# calculate relative abundance (normalize)&lt;br /&gt;
df_phylum &amp;lt;- df_phylum %&amp;gt;%&lt;br /&gt;
  group_by(sample) %&amp;gt;% &lt;br /&gt;
  mutate(rel_abundance = abundance / sum(abundance)) &lt;br /&gt;
&lt;br /&gt;
# plot &lt;br /&gt;
df_phylum %&amp;gt;%&lt;br /&gt;
  ggplot(aes(x=sample, y=rel_abundance, fill=phylum)) +&lt;br /&gt;
  geom_bar(stat=&amp;quot;identity&amp;quot;) +&lt;br /&gt;
  theme(axis.text.x=element_text(angle=90)) +&lt;br /&gt;
  ylab(&amp;quot;Relative abundance&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
##### ALPHA DIVERSITY (diversity within a single sample) #####&lt;br /&gt;
&lt;br /&gt;
#start prepping data by getting sum of species counts across rows&lt;br /&gt;
shannon &amp;lt;- df2 %&amp;gt;%&lt;br /&gt;
  select(-c(phylum, genus)) %&amp;gt;% &lt;br /&gt;
  group_by(species) %&amp;gt;%&lt;br /&gt;
  summarise(across(starts_with(&amp;quot;barcode&amp;quot;), sum), .groups=&amp;quot;drop&amp;quot;) &lt;br /&gt;
&lt;br /&gt;
# convert to a matrix for computations &lt;br /&gt;
mat_species &amp;lt;- shannon %&amp;gt;%&lt;br /&gt;
  select(-1)&lt;br /&gt;
rownames(mat_species) &amp;lt;- shannon$species # set row names as species&lt;br /&gt;
mat_species &amp;lt;- as.matrix(mat_species[,-1]) # drop first col since we have it as row names&lt;br /&gt;
&lt;br /&gt;
# calculate relative abundance with operation applied on columns &lt;br /&gt;
# this divides every element in a column by the total sum of that column arg1-data, arg2-col or row, arg3-metric, arg4-operation; each column sums to 1 &lt;br /&gt;
rel_abundance &amp;lt;- sweep(mat_species, 2, colSums(mat_species), &#039;/&#039;)&lt;br /&gt;
&lt;br /&gt;
# calculate shannon index with custom function applied on columns &lt;br /&gt;
shannon &amp;lt;- apply(rel_abundance, 2, function(p) -sum(p * log(p + 1e-10)))&lt;br /&gt;
&lt;br /&gt;
# prepare final df for platting&lt;br /&gt;
shannon_df &amp;lt;- data.frame(sample = names(shannon), &lt;br /&gt;
                         shannon = shannon)&lt;br /&gt;
&lt;br /&gt;
# Plot&lt;br /&gt;
shannon_df %&amp;gt;% &lt;br /&gt;
  ggplot(aes(x = sample, y = shannon, fill = sample)) +&lt;br /&gt;
  geom_col() +&lt;br /&gt;
  labs(x = &amp;quot;Sample&amp;quot;, y = &amp;quot;Shannon Diversity&amp;quot;, title = &amp;quot;Alpha Diversity per Sample&amp;quot;) +&lt;br /&gt;
  theme_minimal() +&lt;br /&gt;
  theme(axis.text.x = element_text(angle = 45, hjust = 1), legend.position = &#039;none&#039;)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
##### BETA DIVERSITY (Comparing diversity in different samples) #####&lt;br /&gt;
&lt;br /&gt;
# create a dissimilarity matrix between samples using vegdist from vegan package&lt;br /&gt;
# transpose b/c we need rows = samples, cols = species&lt;br /&gt;
# result is pairwise dissimilarities for all samples &lt;br /&gt;
dist_bc &amp;lt;- vegdist(t(rel_abundance), method=&amp;quot;bray&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
# PCoA analysis &lt;br /&gt;
# k=2 dimensions (PC1 and PC2); we&#039;re assigning each sample to a space based on 2 coordinates&lt;br /&gt;
PCoA &amp;lt;- cmdscale(dist_bc, k=2)&lt;br /&gt;
&lt;br /&gt;
# put coordinates into a df for plotting&lt;br /&gt;
PCoA_df &amp;lt;- data.frame(sample = rownames(PCoA),&lt;br /&gt;
                      PC1 = PCoA[,1],&lt;br /&gt;
                      PC2 = PCoA[,2])&lt;br /&gt;
&lt;br /&gt;
# plot&lt;br /&gt;
PCoA_df %&amp;gt;%&lt;br /&gt;
  ggplot(aes(x=PC1, y=PC2, color = sample, label=sample)) +&lt;br /&gt;
  geom_point(size=3) +&lt;br /&gt;
  geom_text(vjust=-0.5) +&lt;br /&gt;
  labs(title = &#039;Beta Diversity&#039;) +&lt;br /&gt;
  xlab(&amp;quot;PCoA1&amp;quot;) + &lt;br /&gt;
  ylab(&amp;quot;PCoA2&amp;quot;)&lt;br /&gt;
  theme_bw() +&lt;br /&gt;
  theme(legend.position = &#039;none&#039;)&lt;br /&gt;
&lt;br /&gt;
# pairwise dissimilarity viz &lt;br /&gt;
# we want a normal matrix, not that vegan object&lt;br /&gt;
bc_matrix &amp;lt;- as.matrix(dist_bc)&lt;br /&gt;
&lt;br /&gt;
# make the heatmap&lt;br /&gt;
pheatmap(bc_matrix, &lt;br /&gt;
         clustering_distance_rows = &amp;quot;euclidean&amp;quot;,&lt;br /&gt;
         clustering_distance_cols = &amp;quot;euclidean&amp;quot;,&lt;br /&gt;
         main = &amp;quot;Pairwise Bray–Curtis Dissimilarity&amp;quot;)&lt;br /&gt;
       &lt;br /&gt;
##### GENUS LEVEL COMPARISONS ACROSS SAMPLES HEATMAP #####&lt;br /&gt;
# for each genus, get total read counts per sample &lt;br /&gt;
df_genus &amp;lt;- df2 %&amp;gt;%&lt;br /&gt;
  select(-c(phylum, species)) %&amp;gt;%&lt;br /&gt;
  group_by(genus) %&amp;gt;%&lt;br /&gt;
  summarise(across(starts_with(&amp;quot;barcode&amp;quot;), sum), .groups=&amp;quot;drop&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
# grab top n genera for plot &lt;br /&gt;
top_genera &amp;lt;- df_genus %&amp;gt;%&lt;br /&gt;
  mutate(total=rowSums(across(starts_with(&amp;quot;barcode&amp;quot;)))) %&amp;gt;% # get row sums across all samples &lt;br /&gt;
  arrange(desc(total)) %&amp;gt;% # most to least&lt;br /&gt;
  slice(1:25) %&amp;gt;% # top n genera &lt;br /&gt;
  select(-total) # remove totals column &lt;br /&gt;
&lt;br /&gt;
genus_mat &amp;lt;- as.matrix(top_genera[,-1]) # convert to matrix, drop the genus col&lt;br /&gt;
rownames(genus_mat) &amp;lt;- top_genera$genus # make genus column row names of the matrix&lt;br /&gt;
&lt;br /&gt;
pheatmap(genus_mat, &lt;br /&gt;
         scale=&amp;quot;row&amp;quot;, # scale the rows to compare across samples &lt;br /&gt;
         clustering_distance_cols=&amp;quot;euclidean&amp;quot;,&lt;br /&gt;
         clustering_method=&amp;quot;ward.D2&amp;quot;,&lt;br /&gt;
         cellwidth = 12,&lt;br /&gt;
         cellheight = 8)&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Main_Page&amp;diff=6703</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Main_Page&amp;diff=6703"/>
		<updated>2026-07-02T10:36:24Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Course/Lecture syllabi */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{|&lt;br /&gt;
|-&lt;br /&gt;
| &lt;br /&gt;
[[File:Protrait-July-2024-Qiu.jpg| x300 px | thumb |&lt;br /&gt;
&amp;lt;strong&amp;gt;Welcome to Qiu Lab Wiki @ Hunter&amp;lt;/strong&amp;gt;&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;Weigang Qiu, Ph.D., Professor, [https://biology.hunter.cuny.edu Department of Biological Sciences]&amp;lt;br&amp;gt; &lt;br /&gt;
[https://hunter.cuny.edu Hunter College] of [https://www.cuny.edu City University of New York]&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
Adjunct Faculty, Weill Cornell Medical College, [https://weill.cornell.edu/units/systems-and-computational-biomedicine Department of Systems and Computational Biomedicine]&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
Belfer Research Building, Room 402&amp;lt;br&amp;gt;413 East  69th Street, New York, NY 10021&amp;lt;br&amp;gt;Office: 1-212-896-0445&amp;lt;br&amp;gt;Email: wqiu-at-(hunter.cuny.edu) &amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
[https://goo.gl/maps/xv1KmaW3XEnxaY1V7 Directions by Google Map]]]&lt;br /&gt;
|&lt;br /&gt;
[[File:Book-cover.jpg | x300 px | thumb | A textbook for microbial genomics &amp;amp; evolution (Spring 2026):&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;[https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 Read excerpts &amp;amp; Order from the publisher using Author Discount (25%) Code &amp;quot;C2508&amp;quot;]&amp;lt;br&amp;gt;&amp;lt;br&amp;gt; [https://digitalgenomes.net Companion codebook (work in progress &amp;amp;#x1F60A;)]]]&lt;br /&gt;
|&lt;br /&gt;
__TOC__&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Fieldwork Gallery==&lt;br /&gt;
&amp;lt;gallery heights=&amp;quot;200&amp;quot; perrow=&amp;quot;4&amp;quot; style=&amp;quot;text-align:left&amp;quot; widths=&amp;quot;250&amp;quot; mode=&amp;quot;packed&amp;quot;&amp;gt;&lt;br /&gt;
File:Tick trip rockefeller.jpeg|Apr 2026, Rockefeller State Park, NY&lt;br /&gt;
File:Tick-trip-2023.jpg|Oct 2023, Long Island, NY&lt;br /&gt;
File:Tick-trip-2022.jpg|June 2022, MA&lt;br /&gt;
File:Tick-trip-2021.jpg|June 2021, MA&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&amp;lt;youtube width=&amp;quot;400&amp;quot; height=&amp;quot;300&amp;quot;&amp;gt;Jb4ACK-GjM0&amp;lt;/youtube&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Lab publications==&lt;br /&gt;
===Lyme Genomics, Evolution, &amp;amp; Ecology===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;4&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
&lt;br /&gt;
File:Screenshot_2024-08-15_094908.png | link=https://journals.asm.org/doi/10.1128/mbio.01749-24 | Akther et al. 2024. &amp;quot;Natural selection and recombination at host-interacting lipoprotein loci drive genome diversification of Lyme disease and related bacteria&amp;quot; &#039;&#039;&#039;&#039;&#039;mBio&#039;&#039;&#039;&#039;&#039; 0:e01749-24. Press coverage: [https://www.gc.cuny.edu/news/cuny-graduate-center-biologists-map-dna-lyme-disease-bacteria CUNY Graduate Center News Story]; [https://hunter.cuny.edu/news/hunter-researcher-maps-dna-of-lyme-disease-bacteria/ Hunter press]&lt;br /&gt;
&lt;br /&gt;
File:Applsci-13-11587-g004.png | link=https://doi.org/10.3390/app132011587 | Di, Chong, Brian Sulkow, Weigang Qiu, and Shipeng Sun. 2023. &amp;quot;Effects of Micro-Scale Environmental Factors on the Quantity of Questing Black-Legged Ticks in Suburban New York&amp;quot; &#039;&#039;&#039;&#039;&#039;Applied Sciences&#039;&#039;&#039;&#039;&#039; 13, no. 20: 11587.&lt;br /&gt;
&lt;br /&gt;
File:Spectrum.01743-22-f003.gif | link=https://doi.org/10.1128/spectrum.01743-22  | Li, Di, Zeglis, Qiu (2022). “Evolution of the &#039;&#039;vls&#039;&#039; antigenic variability locus of the Lyme Disease pathogen and development of recombinant monoclonal antibodies targeting conserved VlsE epitopes”. &#039;&#039;&#039;&#039;&#039;Microbial Spectrum&#039;&#039;&#039;&#039;&#039; 10 (5):1-15. &lt;br /&gt;
&lt;br /&gt;
File:Fig7-small.png | link= https://pubmed.ncbi.nlm.nih.gov/34413477 | Di*, Akther*, Bezrucenkovas, Ivanova, Sulkow, Wu, Mneimneh, Gomes-Solecki, Qiu (2021). &amp;quot;Maximum antigen diversification in a lyme bacterial population and evolutionary strategies to overcome pathogen diversity&amp;quot;. &#039;&#039;&#039;&#039;&#039;The ISME Journal&#039;&#039;&#039;&#039;&#039;. 16, 447-464. (*co-first authors) [https://communities.springernature.com/posts/jenner-s-dilemma-and-how-to-win-evolutionary-arms-races-against-microbial-pathogens Blog Post]&lt;br /&gt;
&lt;br /&gt;
File:Ira-fig2.png | link=https://doi.org/10.21775/9781913652616 | Schwartz, Margos, Casjens, Qiu, Eggers (2020). &amp;quot;Multipartite Genome of Lyme Disease &#039;&#039;Borrelia&#039;&#039;: Structure, Variation and Prophages&amp;quot;. &#039;&#039;&#039;&#039;&#039;Current Issues in Molecular Biology&#039;&#039;&#039;&#039;&#039;. 42:409-454. &lt;br /&gt;
&lt;br /&gt;
File:Fig6-1-Barbour.png | link=https://doi.org/10.1002/9781118960608.gbm01525 | Barbour &amp;amp; Qiu (2019). &#039;&#039;Borreliella&#039;&#039;. In &#039;&#039;&#039;&#039;&#039;Bergey&#039;s Manual of Systematics of Archaea and Bacteria&#039;&#039;&#039;&#039;&#039;. John Wiley &amp;amp; Sons, Inc., in association with Bergey&#039;s Manual Trust. &lt;br /&gt;
&lt;br /&gt;
File:Zjm9990961420002.jpeg | link=https://doi.org/10.1128/jcm.00940-18 | Di L, Wan Z, Akther S, Ying CX, Larracuente A, Li L, Di C, Nunez R, Cucura DM, Goddard NL, Krampis K, Qiu WG. (2018). Genotyping and Quantifying Lyme Pathogen Strains by Deep Sequencing of the Outer Surface Protein C (&#039;&#039;ospC&#039;&#039;) Locus. &#039;&#039;&#039;Journal of Clinical Microbiology&#039;&#039;&#039;. 56(11):e00940-18. &lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Evolution &amp;amp; Learning Algorithms===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;3&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File:Pathogens-12-00388-g003.png | link=https://doi.org/10.3390/pathogens12030388 | Ely, Koh, Ho, Hassan, Pham, Qiu (2023). Novelty Search Promotes Antigenic Diversity in Microbial Pathogens. &#039;&#039;&#039;&#039;&#039;Pathogens&#039;&#039;&#039;&#039;&#039; 12:388.&lt;br /&gt;
File:Pone.0224552.g005.png | link=https://doi.org/10.1371/journal.pone.0224552 | Attie, Sulkow, Di, Qiu (2019). Genetic codes optimized as a traveling salesman problem. &#039;&#039;&#039;&#039;&#039;PLoS ONE&#039;&#039;&#039;&#039;&#039; 14(10): e0224552.&lt;br /&gt;
File:Pcbi.1005677.g001.png | link=https://doi.org/10.1371/journal.pcbi.1005677 | Yan, Deforet, Boyle, Rahman, Liang, Okegbe, Dietrich, Qiu, Xavier (2017). Bow-tie signaling in c-di-GMP: Machine learning in a simple biochemical network. &#039;&#039;&#039;&#039;&#039;PLoS Comput Biol&#039;&#039;&#039;&#039;&#039; 13(8): e1005677. &lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Informatics Tool Development===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File: 12859 2018 2074 Fig1 HTML.png | link=https://pubmed.ncbi.nlm.nih.gov/29499649 | Hernandez, Bernstein, Pagan, Vargas, McCaig, Ramrattan, Akther, Larracuente, Di, Vieira, Qiu. (2018). BbWrapper: BioPerl-based sequence and tree utilities for rapid prototyping of bioinformatics pipelines. &#039;&#039;&#039;&#039;&#039;BMC Bioinformatics&#039;&#039;&#039;&#039;&#039; 19(1):76.&lt;br /&gt;
File: 12859 2014 Article 6488 Fig3 HTML.png | link=https://pubmed.ncbi.nlm.nih.gov/24994456 | Di, Pagan, Packer, Martin, Akther, Ramrattan, Mongodin, Fraser, Schutzer, Luft, Casjens and Qiu. (2014). BorreliaBase: a phylogeny-centered browser of &#039;&#039;Borrelia&#039;&#039; genomes. &#039;&#039;&#039;&#039;&#039;BMC Bioinformatics&#039;&#039;&#039;&#039;&#039; 15:233.&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*[http://scholar.google.com/citations?hl=en&amp;amp;user=Ds6u39QAAAAJ) Full list by Google Scholar]&lt;br /&gt;
*[http://www.ncbi.nlm.nih.gov/sites/myncbi/weigang.qiu.1/bibliography/42770924/public/ Full list by NCBI Bibliography]&lt;br /&gt;
last update: March 20, 2023&lt;br /&gt;
&lt;br /&gt;
==Lab members and trainees==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
!Year/Period !!Doctoral members &amp;amp; trainees !!Other members &amp;amp; trainees&lt;br /&gt;
|-&lt;br /&gt;
|Current Academic Year&lt;br /&gt;
(Spring &amp;amp; Summer 2026)&lt;br /&gt;
||&lt;br /&gt;
*Edee Amin: CUNY Graduate Center, Bio/EEB doctoral program&lt;br /&gt;
* Dr Yozen Hernandez: System administrator (part-time), Ph.D. from Boston University&lt;br /&gt;
||&lt;br /&gt;
*Tasmina Hassan: Hunter Bio/CS&lt;br /&gt;
*Mathew DiCicco: Hunter Math/CS&lt;br /&gt;
*Tara Doma Lama: Hunter Bio/Bioinformatics&lt;br /&gt;
*Hagar Abuzaid: Hunter Bio&lt;br /&gt;
|-&lt;br /&gt;
|Alumni (Since Fall 2002)&lt;br /&gt;
||&lt;br /&gt;
* Brandon Ely (2026): CUNY Grad Center, Biology/MCD doctoral program&lt;br /&gt;
* Li Li (Lily, 2023): CUNY Grad Center, Biology/EEB doctoral program&lt;br /&gt;
*Dr Lia Di: Ph.D. Postdoctoral Research Associate, from Nanjing Agricultural University &amp;amp; Wisconsin Blood Institute&lt;br /&gt;
*Dr Saymon Akther (2022): CUNY Grad Center, Biology/EEB&lt;br /&gt;
*Dr Rayees Rahman: Hunter Bio/Bioinformatics, Ph.D. from Mt Sinai Medical School&lt;br /&gt;
*Dr Che Martin (2013): CUNY Grad Center, Biology/MCD&lt;br /&gt;
*Dr James Haven (2011): CUNY Grad Center, Biology/MCD&lt;br /&gt;
*Dr Tika Sukarna (2009): CUNY Grad Center, Biology/MCD&lt;br /&gt;
*Dr Juan Coronado (2008): CUNY Grad Center, Biology/MCD (Dr Peter Lipke)&lt;br /&gt;
*Dr William McCaig: CUNY BA, Ph.D. from Stony Brook University&lt;br /&gt;
*Dr Vincent Xue: CUNY CS/Bioinformatics, Ph.D. from MIT&lt;br /&gt;
*Dr Fubin Li: CUNY Grad Center, Biology/MCD (Dr Laurel Eckhardt)&lt;br /&gt;
*Dr Oliver Attie: Postdoctoral Research Associate, Ph.D. from NYU&lt;br /&gt;
||&lt;br /&gt;
(published coauthors)&amp;lt;br&amp;gt; &lt;br /&gt;
*Brian Sulkow: CUNY Grad Center, Mathematics&lt;br /&gt;
* Winston Koh: Hunter Bio/CS&lt;br /&gt;
*Eamen Ho: Hunter Bio/Bioinformatics&lt;br /&gt;
*Ahn Pham: Hunter Bio/Bioinformatics&lt;br /&gt;
*Chris Panlasigui: Hunter Bio/Bioinformatics&lt;br /&gt;
*Amanda Amanda Larracuente: CUNY Grad Center, Biology/MCD&lt;br /&gt;
* Pedro Pagan: Hunter Bio/Bioinformatics&lt;br /&gt;
*Edgaras Bezrucenkovas: Hunter Chem/Bioinformatics&lt;br /&gt;
*Girish Ramrattan: Hunter Bio/Bioinformatics&lt;br /&gt;
* Levy Vargas: Hunter Bio/Bioinformatics&lt;br /&gt;
*Chong Di: Hunter Geography&lt;br /&gt;
*Roy Nunez: Hunter Bio&lt;br /&gt;
*Mei Wu: CUNY City Tech&lt;br /&gt;
*Desiree Pante: Hunter Bio&lt;br /&gt;
*Saimtun Shipa: Hunter Stat (MA)&lt;br /&gt;
*Bing Wu: Hunter Bio/Biotechnology&lt;br /&gt;
* Svidatoslav Kendall (Slav): Hunter Biology&lt;br /&gt;
*Philip Romov: Hunter CS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Web Apps==&lt;br /&gt;
===Apps with Collaborators (or from published papers)===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;4&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File:Screenshot_2024-06-19_205621.png | link=https://cov.genometracker.org/finalelisa-app/ | Borrelia diagnostic antigens (App developed by Liann Aris-Henry; Data from [https://journals.asm.org/doi/10.1128/jcm.01142-19 Arumugam et al (2019)])&lt;br /&gt;
File:Screenshot_2024-06-18_144931.png | link=https://cov.genometracker.org/borrelia-app/ | Borrelia growth transcriptome (App developed by Laziz Asamov; Data from [https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0164165 Arnold et al (2016)])&lt;br /&gt;
File: Silac.png | link=http://borreliabase.org/~wgqiu/Polotskaia_etal_2014/supp-fig-s1/ | Data &amp;amp; [http://borreliabase.org/~wgqiu/Polotskaia_etal_2014/supp-table-s1/ GSEA] associated with Polotskaia_etal_2014 (with Bargonetti Lab @Hunter)&lt;br /&gt;
File:Screenshot_2023-03-24_122016.png | link=http://borreliabase.org/~wgqiu/clickme-khalikuz/temp-Points.html | Phosphoproteome MDM2KD data set (with Bargonetti Lab @Hunter)&lt;br /&gt;
File:Screenshot_2023-03-24_122312.png | link=http://borreliabase.org/~wgqiu/mpai-v3/ | Genes &amp;amp; Pathways Associated with mTP53 (with Bargonetti Lab @Hunter)&lt;br /&gt;
File:Spombe.png | link=Spombe | S. pombe transcriptomes (with Zhong Lab @Hunter)&lt;br /&gt;
File:Screenshot 2023-03-24 122223.png | link=http://borreliabase.org/~wgqiu/gbs-browser-v3/ | Group B Streptococcus (GBS) genome browser (with Wu Lab @Shenzhen)&lt;br /&gt;
File:Screenshot 2023-03-24 122245.png | link=http://borreliabase.org/~wgqiu/oneKGenome/ | 1K genome (with Ogunwobi Lab @Hunter)&lt;br /&gt;
File:Screenshot 2023-03-24 122450.png | link=http://borreliabase.org/~wgqiu/E_faecalis/ | E faecalis genome browser (with Morales Lab @WCMC)&lt;br /&gt;
File:Screenshot 2023-03-24 122145.png | link=http://borreliabase.org/~wgqiu/carmen-proteomics/ | Proteomics GSEA results (with Melendez Lab @Hunter) &lt;br /&gt;
File:Screenshot 2023-03-24 122051.png | link=http://borreliabase.org/~wgqiu/Tcell-pHis-ACS/ | pHis Protein Browser (with Skolnik lab @NYU)&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Qiu Lab Apps===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;3&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File:Screenshot 2023-03-24 122558.png | link=http://borreliabase.org | Lyme genome browser&lt;br /&gt;
File:Screenshot 2023-03-24 133320.png | link=http://borreliabase.org/~wgqiu/pf-trees/ | Lyme pathogen plasmid partitioning gene trees&lt;br /&gt;
File:Screenshot 2023-03-24 133636.png | link=http://borreliabase.org/vls-finder/ | vls Finder in Lyme pathogen genomes&lt;br /&gt;
File:Screenshot 2023-03-24 122359.png | link=http://borreliabase.org/~wgqiu/code-wheel | Codon Wheel&lt;br /&gt;
File:PhyloView.png | link=http://borreliabase.org/~wgqiu/PhyloView | Co-visualization of a tree with an alignment and character matrix&lt;br /&gt;
File:Screenshot 2023-03-24 122521.png | link=http://borreliabase.org/~wgqiu/ospC-sequencing | OspC amplicon sequencing from ticks&lt;br /&gt;
File:Screenshot 2023-03-24 122340.png | link=http://borreliabase.org/~wgqiu/asm-challenge | Genomic epidemiology of a Salmonella outbreak (ASM Challenge)&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Curricular Development &amp;amp; Bioinformatics/QuBi Advising ==&lt;br /&gt;
*QuBi advisors: Weigang Qiu, Ntino Krampis, Rabindra Mandal (Biology); Saad Mneimeih, Lei Xie (CS); Akira Kawamura (Chem); Dana Sylvan (Math &amp;amp; Stats) &lt;br /&gt;
**Permission for non-Biology majors to take BIOL203 &amp;amp; BIOL425, every Spring&lt;br /&gt;
**Collect names,  major, and IDs to send to course coordinator to grant permission. Waive BIOL10200 pre-reqs for taking BIOL203.&lt;br /&gt;
*Curricular resources:&lt;br /&gt;
**[http://biology.hunter.cuny.edu/index.php?option=com_content&amp;amp;view=article&amp;amp;id=66&amp;amp;Itemid=73 Biology courses and pre-reqs]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/BIO1-BA Hunter Biology Major 1 (including the Bioinformatics Option)]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/CHEM2-BA Hunter Chemistry Major 2 (including the Bioinformatics Option)]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/COMPSCI-BA Hunter Computer Science (including Bioinformatics Concentration)]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/MATH-BA Hunter Mathematics (including the Quantitative Biology Concentration)]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/STATS-BA Hunter Statistics (including the Quantitative Biology Concentration)]&lt;br /&gt;
*QuBi advising:&lt;br /&gt;
**Declaration of Bioinformatics concentration: In-person advising to work out the semester-by-semester courses&lt;br /&gt;
**Approve on department spreadsheet (or send email to  &amp;quot;Samantha Sheppard-Lahiji&amp;quot; and &amp;quot;HTR Bio&amp;quot; &amp;lt;biology@hunter.cuny.edu&amp;gt;)&lt;br /&gt;
**Students should take Bioinformatics-specific electives (8 cred; see Hunter Catalog below), &#039;&#039;&#039;not general electives&#039;&#039;&#039;&lt;br /&gt;
***Examples: Anthrop302 (3 cr); Chem333 (3 cr); BIOL47119 &amp;amp; BIOL47120 (3cr); BIOL48002 (2 cr)&lt;br /&gt;
**Students need to take BIOL48002 (2 cr), which counts towards as research credit, to graduate as honors&lt;br /&gt;
*General advising:&lt;br /&gt;
**~40 students every semester. Send out emails to students. Go through student courses by Email or by appointment&lt;br /&gt;
**Recommend new math courses: &#039;&#039;&#039;MATH15200 &amp;amp; STAT21350&#039;&#039;&#039;&lt;br /&gt;
*Hosting QuBi students in lab&lt;br /&gt;
**This is to enhance the informatics and coding skills of our students&lt;br /&gt;
**Students should register and get &#039;&#039;&#039;BIOL48002&#039;&#039;&#039; credits, which counts towards their elective credits &amp;amp; eligibility for honors&lt;br /&gt;
**3-5 students per semester&lt;br /&gt;
*Outside research opportunities&lt;br /&gt;
**MIT Quantitative Workshop (first week of January, in Boston). Coordination with CS (Saad &amp;amp; Susan Epstein) in Fall&lt;br /&gt;
**Simons Foundation/Flatiron Institute Center for Computational Biology (CCB) Internship program. Open House in Spring&lt;br /&gt;
&lt;br /&gt;
==Course/Lecture syllabi==&lt;br /&gt;
*[[NYRaMP-Informatics-2025|NYRaMP Workshop (August 2025, by Brandon Ely)]]&lt;br /&gt;
*[[Computational Genomics Summer 2026 | KIZ Genomics Workshop, Summer 2026]]&lt;br /&gt;
*[[Computational Genomics (KIZ, Fall 2024)]]&lt;br /&gt;
*[[NYRaMP-Informatics-2024|NYRaMP Workshop (August 2024)]]&lt;br /&gt;
*BIOL47120 BioMedical Genomics (Spring 2024). Tutorials: [https://borreliabase.org/~wgqiu/tutorial-markdown.html R Markdown] [https://borreliabase.org/~wgqiu/cluster-analysis.html Cluster analysis] [https://borreliabase.org/~wgqiu/scRNA-analysis.html single-cell RNA-seq]&lt;br /&gt;
*BIOL425 Computational Molecular Biology (Spring, 2023). [https://github.com/weigangq/CSB-BIOL425/tree/master/lecture-materials Lecture material on github]&lt;br /&gt;
*BIOL714 Cell Biology: [http://borreliabase.org/~wgqiu/r-demo-2024.html R Demo (Spring 2024)] [http://borreliabase.org/~wgqiu/r-demo-2023.html R Demo (Spring 2023)]&lt;br /&gt;
*QuBi module: [[QuBi/module/bio203-lab12—2022|BIOL20300 Molecular Genetics, Lab 12 (2023)]]&lt;br /&gt;
*QuBi module: [[QuBi/modules/biol203-geno-pheno-association-2022|BIOL20300 Molecular Genetics, Lab 13 (2022)]]&lt;br /&gt;
*QuBi module: [[QuBi/modules/biol303|BIOL30300 Cell Biology, Bioinformatics Lab (transcriptome analysis)]]&lt;br /&gt;
*[[BigData 2020|Big Data (Summer, 2020)]]&lt;br /&gt;
*[[BioMed-R-2020|BIOL47120 Biomedical Genomics II (Spring, 2020)]] &lt;br /&gt;
**[http://borreliabase.org/~wgqiu/tutorial-markdown.html Tutorial: R Markdown (Spring 2024)]&lt;br /&gt;
**[http://borreliabase.org/~wgqiu/cluster-analysis.html Tutorial: Cluster analysis (Spring 2024)]&lt;br /&gt;
**[http://borreliabase.org/~wgqiu/scRNA-analysis.html Tutorial: single-cel transcriptome analysis (Spring 2024)]&lt;br /&gt;
*[[Biol425 2020|BIOL425 Computational Molecular Biology (Spring, 2020)]]&lt;br /&gt;
*[[Biol375 2019|BIOL37500, Molecular Evolution (Fall, 2019)]]&lt;br /&gt;
*[[Southwest-University|Southwest University R course (Summer, 2019)]]&lt;br /&gt;
*[[Biol20N02 2017|Analysis of Biological Data (Spring, 2017)]]&lt;br /&gt;
*[[Bioinformatics_Workshop_2014|Bioinformatics Workshop (Summer, 2014)]]&lt;br /&gt;
&lt;br /&gt;
==SARS-CoV-2 genome evolution==&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;3&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File:Cov-fig1.jpg | Akther, Bezrucenkovas, Sulkow, Panlasigui, Qiu, Di (April, 2020). &amp;quot;CoV Genome Tracker: tracing genomic footprints of Covid-19 pandemic&amp;quot;. &#039;&#039;&#039;&#039;&#039;[https://www.biorxiv.org/content/biorxiv/early/2020/04/14/2020.04.10.036343.full.pdf BioRxiv]&#039;&#039;&#039;&#039;&#039;; [https://cov.genometracker.org/ Web app: SARS-CoV-2 Genome Tracker]; Github: https://github.com/weigangq/cov-browser&lt;br /&gt;
File:Rec-fig2.png | Akther, Li, Martin, Di, Sulkow, Pante, Bezrucenlovas, Luft, Qiu (May, 2020). &amp;quot;Origin, recombination, and missed opprotunities:  a genomic perspective of the first 100 days of COVID-19 pandemic&amp;quot;. (Unpublished).&lt;br /&gt;
File:Cov-fig3-trace.png | Akther, Bezrucenlovas, Li, Sulkow, Di, Pante, Martin, Luft, Qiu (Sep, 2021). &amp;quot;Following the Trail of One Million Genomes: Footprints of SARS-CoV-2 Adaptation to Humans&amp;quot;. &#039;&#039;&#039;&#039;&#039;[https://www.biorxiv.org/content/biorxiv/early/2021/05/10/2021.05.07.443114.full.pdf BioRxiv link]&#039;&#039;&#039;&#039;&#039;: . Github:  https://github.com/weigangq/cov-db&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Lab Resources &amp;amp; Protocols==&lt;br /&gt;
*OneRF usage (Unofficial; login needed): https://sites.google.com/view/onerf/home&lt;br /&gt;
*[https://runestone.academy/ns/books/published/thinkcspy/index.html How to think like a computer scientist: An interactive Python programming book]&lt;br /&gt;
* Nanopore sequencing protocols&lt;br /&gt;
**DNA barcoding: https://nanopore4edu.org/latest/annotated_experiments/dna_barcoding/&lt;br /&gt;
**Yeast genomes: https://nanoporetech.com/document/extraction-method/yeast-dna&lt;br /&gt;
*[[Monte Carlo Club]]&lt;br /&gt;
*[[NY-RaMP Mentoring]]&lt;br /&gt;
*Borreliella genome sequencing consortium: Weekly meetings (Tu @11): Since Jan 2023&lt;br /&gt;
*Borreliella diagnostic antigens (Fall 2023-Fall 2027):&lt;br /&gt;
**Zoom call (Jan 23, 2024)&lt;br /&gt;
** Next meeting: March 23, 2024&lt;br /&gt;
*Qiu lab network [[First Time Guide|first-time user guide]]&lt;br /&gt;
*Qiu lab Github repositories: https://github.com/weigangq/?tab=repositories&lt;br /&gt;
*[[Mini-Tutorals|Mini-Protocols]] (frequently used computer codes and pipelines)&lt;br /&gt;
*Python tutorial: https://wiki.genometracker.org/~weigang/Intro_to_Python.html&lt;br /&gt;
* [[Tick protocol|ick handling protocols]]&lt;br /&gt;
*[[A Primer on the Cluster System at Hunter|Hunter HPC Usage]]&lt;br /&gt;
*[https://r4ds.hadley.nz/ R for Data Science (2e)], (2024) by Wickham, Grolemund &amp;amp; Çetinkaya-Rundel ([https://bookdown.org/ Bookdown version])&lt;br /&gt;
*Borrelia Genome Consortium:  https://www.ncbi.nlm.nih.gov/bioproject/PRJNA431102/&lt;br /&gt;
*Canadian Bbsl genome assemblies: https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA1130942&lt;br /&gt;
* Nanopore sequencing resources:&lt;br /&gt;
**eBook: https://store.nanoporetech.com/us/minion.html&lt;br /&gt;
&lt;br /&gt;
==Wiki Help==&lt;br /&gt;
*[https://www.mediawiki.org/wiki/Special:MyLanguage/Manual:Configuration_settings Configuration settings list]&lt;br /&gt;
*[https://www.mediawiki.org/wiki/Special:MyLanguage/Manual:FAQ MediaWiki FAQ]&lt;br /&gt;
*[https://lists.wikimedia.org/postorius/lists/mediawiki-announce.lists.wikimedia.org/ MediaWiki release mailing list]&lt;br /&gt;
*[https://www.mediawiki.org/wiki/Special:MyLanguage/Localisation#Translation_resources Localise MediaWiki for your language]&lt;br /&gt;
*[https://www.mediawiki.org/wiki/Special:MyLanguage/Manual:Combating_spam Learn how to combat spam on your wiki]&lt;br /&gt;
*Consult the [[mediawikiwiki:Special:MyLanguage/Help:Contents|User&#039;s Guide]] for information on using the wiki software.&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6702</id>
		<title>Computational Genomics Summer 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Computational_Genomics_Summer_2026&amp;diff=6702"/>
		<updated>2026-07-02T10:31:02Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: Created page with &amp;quot;center &amp;lt;center&amp;gt;9-12 noon, July 13 - Aug 6, 2026&amp;lt;/center&amp;gt; &amp;lt;center&amp;gt;&amp;#039;&amp;#039;&amp;#039;Guest Instructor:&amp;#039;&amp;#039;&amp;#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt; &amp;lt;center&amp;gt;&amp;#039;&amp;#039;&amp;#039;Office:&amp;#039;&amp;#039;&amp;#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/cente...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Banner-comp-genomics.png|800px|center]]&lt;br /&gt;
&amp;lt;center&amp;gt;9-12 noon, July 13 - Aug 6, 2026&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Guest Instructor:&#039;&#039;&#039; Weigang Qiu, Ph.D.&amp;lt;br&amp;gt;Professor, Department of Biological Sciences, City University of New York, Hunter College &amp;amp; Graduate Center&amp;lt;br&amp;gt;Adjunct Faculty, Department of Systems and Computational Biomedicine, Weil Cornell Medical College&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Office:&#039;&#039;&#039; B402 Belfer Research Building, 413 East 69th Street, New York, NY 10021, USA&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Email:&#039;&#039;&#039; wqiu@hunter.cuny.edu&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Lab Website:&#039;&#039;&#039; https://wiki.genometracker.org&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;&#039;&#039;&#039;Host &amp;amp; Assistants&#039;&#039;&#039;&amp;lt;/center&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
==Course Overview==&lt;br /&gt;
Welcome to Computational Genomics, a 9-week computer workshop for graduate students. A genome is the total genetic content of an organism. Driven by breakthroughs such as the decoding of the first human genome and next-generation DNA -sequencing technologies, biomedical sciences are undergoing a rapid and irreversible transformation into a highly data-intensive field. &lt;br /&gt;
&lt;br /&gt;
Genome information is revolutionizing virtually all aspects of life sciences including basic research, medicine, and agriculture. Meanwhile, use of genomic data requires life scientists to be familiar with concepts and skills in biology, computer science, as well as data analysis. &lt;br /&gt;
&lt;br /&gt;
This workshop is designed to introduce computational analysis of genomic data through hands-on computational exercises, using published studies.&lt;br /&gt;
&lt;br /&gt;
The pre-requisites of the course are college-level courses in molecular biology, cell biology, and genetics. Introductory courses in computer programming and statistics are preferred but not strictly required.&lt;br /&gt;
&lt;br /&gt;
==Learning goals==&lt;br /&gt;
By the end of this course successful students will be able to: &lt;br /&gt;
* Use Linux commands &amp;amp; compose simple shell scripts to automate a bioinformatics pipeline&lt;br /&gt;
* Program in Python for parsing texts and simulating evolution &lt;br /&gt;
* Visualize data and perform statistical analysis using R/RStudio&lt;br /&gt;
* Compose a bioinformatics research report&lt;br /&gt;
&lt;br /&gt;
==Web Links==&lt;br /&gt;
* Install R base: https://cloud.r-project.org&lt;br /&gt;
* Install R Studio (Desktop version): http://www.rstudio.com/download&lt;br /&gt;
* Download: [http://www.r4all.org/books/datasets R datasets]&lt;br /&gt;
* A reference book: [https://r4ds.hadley.nz/ R for Data Science (Wickharm et al)]&lt;br /&gt;
* Github repository: [https://github.com/weigangq/CSB-BIOL425/tree/master/lecture-materials Computational Skills for Biologists (Allesina &amp;amp; Wilmes)]&lt;br /&gt;
&lt;br /&gt;
==Assignments, Quizzes, and Final Report==&lt;br /&gt;
Student performance will be evaluated by attendance, three (3) quizzes and a final presentation&lt;br /&gt;
* Attendance &amp;amp; participation: 60 pts&lt;br /&gt;
* Open-Book Quizzes: 3 x 30 pts = 60 pts&lt;br /&gt;
* Final presentation: 30 pts&lt;br /&gt;
Total: 150 pts&lt;br /&gt;
&lt;br /&gt;
==Course Schedule==&lt;br /&gt;
===Week 1===&lt;br /&gt;
====Session 1, 7/13 (Monday)====&lt;br /&gt;
* Introudction. Lecture slides: [[File:QiuLab-CUNY-Hunter.pdf|thumb]]&lt;br /&gt;
* Computer setup: gitee accounts (for course management). [https://edu.gitee.com/huntercollege/courses/3030/noticeboard Course link on gitee]&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 1: Genomics &amp;amp; Data Science&amp;lt;/span&amp;gt; PDF file: [[File:KIZ-survey-1.pdf|thumb]]&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: blue&amp;quot;&amp;gt;Survey 2: Tree-thinking Skills&amp;lt;/span&amp;gt; PDF file: [[File:Pretest-1.pdf|thumb]]&lt;br /&gt;
* Computer setup: Linux accounts (on the &amp;quot;phylonet.net&amp;quot; server)&lt;br /&gt;
* Lecture: Tree-thinking Skills. Lecture slides: [[File:Phylogeny-lecture-slides-2024.pdf]]&lt;br /&gt;
&lt;br /&gt;
====Session 2, 7/14 (Tuesday)====&lt;br /&gt;
* Git usages (by Mr Liu)&lt;br /&gt;
** Demo: homework submission using Gitee&lt;br /&gt;
** Download the course repository: &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/comp-genomics-kiz.git&amp;lt;/code&amp;gt; &lt;br /&gt;
* Linux Tutorial I. [[File:Intro-unix-KIZ.pdf|thumb]]&lt;br /&gt;
* Phylogenetics lecture: [[File:Part-1-tree-thinking.pdf|thumb]]&lt;br /&gt;
** Tree manipulations: reroot &amp;amp; tree distances&lt;br /&gt;
** Gene tree vs Species tree; orthologous and paralogous genes; tree pruning &amp;amp; collapsing&lt;br /&gt;
&lt;br /&gt;
====Session 3, 7/16 (Thursday)====&lt;br /&gt;
* Review: Tree manipulations&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #1. tree terms &amp;amp; tree manipulations&amp;lt;/span&amp;gt; (20 pts); open-book; 9-10am, in-class&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: green&amp;quot;&amp;gt;Linux Tutorial II: BpWrapper Toolkit (https://github.com/bioperl/p5-bpwrapper)&amp;lt;/span&amp;gt; Updated sldies: [[File:Intro-unix-KIZ.pdf|thumb]]&lt;br /&gt;
&lt;br /&gt;
===Week 4, Thursday, Oct 31, 2024 (Halloween)===&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: red&amp;quot;&amp;gt;Quiz #2. Linux commands &amp;amp; BpWrappper toolkits&amp;lt;/span&amp;gt;: 20 pts, 9-10:30am&lt;br /&gt;
* &amp;lt;span style=&amp;quot;color: orange&amp;quot;&amp;gt;AFSV Genomics I &amp;lt;/span&amp;gt;: Download and align genomes. &lt;br /&gt;
** Fork a copy of [https://edu.gitee.com/huntercollege/projects/696311/repos/huntercollege/afsv-genomics/sources the ASFV project repository]. &lt;br /&gt;
** Instructions to fork a repository (prepared by Mr Liu): https://zwmqn249t3y.feishu.cn/wiki/OVdEwB00ciNx9Xkxvxsc7lJEncc?from=from_copylink&lt;br /&gt;
** Project overview: Lecture slides: [[File:Afsv-project-kiz.pdf|thumb]]&lt;br /&gt;
&lt;br /&gt;
===Week 5, Thursday, Nov 7, 2024===&lt;br /&gt;
* ASFV project: &lt;br /&gt;
** &amp;lt;code&amp;gt; git clone https://gitee.com/huntercollege/asfv-genomics.git &amp;lt;/code&amp;gt;&lt;br /&gt;
** Align genomes: protocol-1&lt;br /&gt;
** SNP calls; quality check by Ts/Tv ratio (Protocol 1)&lt;br /&gt;
** Annotate SNPs: Protocol 2&lt;br /&gt;
** IQ-tree to get genome tree: Protocol-3&lt;br /&gt;
* Trait evolution and comparative analysis [[File:Trait-evolution-KIZ.pdf|thumb]]&lt;br /&gt;
* Assignment: Pre-test 2&lt;br /&gt;
&lt;br /&gt;
===Week 6, Thursday, Nov 14, 2024===&lt;br /&gt;
* Computer setup: R &amp;amp; RStudio; &amp;lt;span style=&amp;quot;color: orange&amp;quot;&amp;gt;R Tutorial Part 1: Basic R &amp;amp; Data manipulations&amp;lt;/span&amp;gt;. Lecture slides [[File:R-tutorials-KIZ-part-1.pdf|thumb]]&lt;br /&gt;
* Assignment: Practice-1 &amp;amp; -2&lt;br /&gt;
&lt;br /&gt;
===Week 7, Thursday, Nov 21, 2024===&lt;br /&gt;
* R Tutorial 2. Data visualization &amp;amp; statistics. Slides: [[File:R-tutorials-KIZ-part-2.pdf|thumb]]&lt;br /&gt;
** Assignment: Practice-3 &amp;amp; -4&lt;br /&gt;
* Final project:&lt;br /&gt;
** Introduction to Ka/Ks analysis&lt;br /&gt;
** File distribution: each student is assigned 10 random genes; Run &amp;lt;code&amp;gt;git pull&amp;lt;/code&amp;gt;; or if you haven&#039;t cloned the repository, run &amp;lt;code&amp;gt;git clone https://gitee.com/huntercollege/asfv-genomics.git&amp;lt;/code&amp;gt;&lt;br /&gt;
** Follow the protocol in &amp;quot;doc/protocol-4-paml.txt&amp;quot;&lt;br /&gt;
&lt;br /&gt;
===Week 8, Thursday, Nov 28, 2024 (Thanksgiving)===&lt;br /&gt;
* R Tutorial 3. [https://borreliabase.org/~wgqiu/r-demo-2024.html Cluster analysis]: &lt;br /&gt;
** Part 4. Heatmap (hierarchical clustering) &amp;amp; principal component analysis (PCA)&lt;br /&gt;
** Part 5. Gene expression analysis. &lt;br /&gt;
** Assignment: reproduce the cluster analysis&lt;br /&gt;
* Final project: [https://borreliabase.org/~wgqiu/asfv-genomics-visual.html R Markdown Demo]&lt;br /&gt;
** Visualize tree with &amp;lt;code&amp;gt;ggtree&amp;lt;/code&amp;gt;&lt;br /&gt;
** Plot Ka/Ks for genes&lt;br /&gt;
** Run IQ-TREE to obtain site-specific rates; Plot site-specific rates&lt;br /&gt;
&lt;br /&gt;
===Week 9, Thursday, Dec 5, 2024===&lt;br /&gt;
* Final presentations: student demos &amp;amp; trouble shooting&lt;br /&gt;
&lt;br /&gt;
===Week 10, Thursday, Dec 10, 2024===&lt;br /&gt;
* ASFV project overview slides: [[File:Asfv-project-kiz-Dec-5-2024.pdf|thumb]]&lt;br /&gt;
* Final presentations (30 pts):&lt;br /&gt;
** (10 pts) 3 slides &amp;amp; 5 min&lt;br /&gt;
** (15 pts) Show &amp;amp; interpret results for a single gene (no need to show all genes on PPT, although you will upload all MLC and tree files). Pick a genes that is the most interesting to you (e.g., significant position selection, apparent inconsistency with the genome tree, important gene function, etc).&lt;br /&gt;
** (5 pts) Conclusions &amp;amp; future directions. Find gene name and function from this paper: https://www.mdpi.com/2076-2615/14/15/2187 (you need to first find the gene name using the &amp;quot;BA71.gff3&amp;quot; file with &amp;lt;code&amp;gt;grep&amp;lt;/code&amp;gt;) &lt;br /&gt;
* Course evaluation&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=A_Primer_on_the_Cluster_System_at_Hunter&amp;diff=6701</id>
		<title>A Primer on the Cluster System at Hunter</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=A_Primer_on_the_Cluster_System_at_Hunter&amp;diff=6701"/>
		<updated>2026-06-23T17:07:29Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: /* Configuring bioconda */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
=What is a Cluster System?=&lt;br /&gt;
[[File:Ccs-overview.png|thumb|frameless|&#039;&#039;&#039;Figure 1&#039;&#039;&#039; &amp;lt;sub&amp;gt;The general idea of a cluster system. There is a login node that is directly connected to the other nodes of the cluster&amp;lt;/sub&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
A &#039;&#039;&#039;Cluster System&#039;&#039;&#039; is a set of connected computers that work together in order to perform tasks, unlike many of the portals at hunter that require you to first login to a head node, and then to a working node. &lt;br /&gt;
*For example to enter the Qiu lab servers you must first login to Darwin.hunter.cuny.edu and then to a compute node such as Wallace.hunter.cuny.edu &lt;br /&gt;
On a cluster system all of the nodes connected can be viewed as a single system. Under this system all you would need to do is login to the head node, and from there you can run your programs. However,  unlike the non-cluster computer servers at hunter, in order to fully utilize the cluster you must submit your job to a job scheduler that controls which node to run the job on. &lt;br /&gt;
&lt;br /&gt;
=Basics of Using a Cluster System=&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
*What programs are currently installed on the cluster System? &lt;br /&gt;
The normal way of running programs on a linux machine would be to type the name of the program into the terminal, add any options to the command, and press enter. For example&amp;quot; &lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
$&amp;gt;blastp -query test.fas -db nr -remote -outfmt 6 &lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
However this will not work in a cluster. &lt;br /&gt;
In order to run programs on a cluster, the cluster must first have those programs installed (or you may have to compile the program and run it locally) and then you must load the program on to your terminal.&lt;br /&gt;
To check the available programs on the Hunter College cluster use this command: &lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
module avail&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
This command shows what modules are currently available to use. To save the list of available programs use this command:&lt;br /&gt;
 &amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
module avail 2&amp;gt; ~/available_apps.txt &lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
This command stores the present list programs installed on the cluster to a file called &amp;quot;available_apps.txt&amp;quot; to your home directory. &lt;br /&gt;
&#039;&#039;&#039;This file will only show the present programs installed on the cluster, you will need to update this file in the future to show new applications!&#039;&#039;&#039;&lt;br /&gt;
*How do we load a program on the cluster? &lt;br /&gt;
Now that we know what are the available apps on the cluster, we need to load the application into our terminal. To do so we use the command: &lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
module load &amp;lt;Program Name&amp;gt;&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt; &lt;br /&gt;
For example, to load blast on the cluster we need to run the command: &lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
module load R/3.2.2&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
Notice how we have a version number when we run the module load command? &lt;br /&gt;
This means that on the cluster we can use different versions of a program! This is useful for those situations when an updated version of a program breaks compatibility with some code you&#039;ve written or some modules you&#039;ve downloaded, or if you like one version of a program more than the other.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
In order to use a different version of a program we first need to unload our program then load up the correct version of the program: &lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
module unload R/3.2.2&lt;br /&gt;
module load R/3.1.0&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*How do we run a program on the cluster?&lt;br /&gt;
Now lets go back to our previous blast example. Lets first load up the module: &lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
module load blast/2.2.31&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
Notice now, when you tab &amp;quot;bl&amp;quot; on the command line all programs associated with blast comes up, such as blastn / blastp, etc. &lt;br /&gt;
Now you may be tempted to run blast, or any program, the normal way, that is:&lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
blastp -query test.fas -db nr -remote -outfmt 6&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;Don&#039;t do this!!!&#039;&#039;&#039; When you run the command like this, you are running it on the head node!! &lt;br /&gt;
You aren&#039;t using the power of the cluster to do the work, that and you may be clogging up the little computational power (relative to the compute nodes) the head node has and making everyone else experience on the cluster laggy and unstable!&lt;br /&gt;
To run a job on the cluster you must run it through a scheduler, in our case the cluster uses SLURM, but the same principals apply to other cluster formats such as LSF or BSUB. &lt;br /&gt;
&lt;br /&gt;
 &lt;br /&gt;
&#039;&#039;&#039;The correct way of submitting a job to the cluster from the terminal is to use this command&#039;&#039;&#039;: &lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
srun &amp;lt;command&amp;gt;&amp;gt; &lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt; &lt;br /&gt;
This will submit the command to a compute node, and run the job interactively, that is, all output will redirect to the terminal you are using, for example:&lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
module load cdhit/4.64&lt;br /&gt;
srun cd-hit -h&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
Will output the help directly to your terminal. While this is fine for quick and dirty programs like cdhit, when you run blast you&#039;ll get a blank output until the blast is complete. Try it for your self: &lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
blastp -query test.fas -db nr -remote -outfmt 6&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
If you close the terminal, or if you lose your connection to the internet during this time, all progress will be lost. So how do we submit jobs to the cluster and just let things run without worrying about our job dying? Also can we allocate specific resources to our programs, like the number of cpu&#039;s it needs, or the amount RAM? To do this we need to write some scripts, so open up a text editor and head to the next section, yo.&lt;br /&gt;
&lt;br /&gt;
==Cluster Scripting==&lt;br /&gt;
In order to submit a job to the cluster and have it run without keeping a terminal open &#039;&#039;&#039;we need to write a script&#039;&#039;&#039;. &lt;br /&gt;
You can use this template script as a basis for any future cluster scripts you may need to write: &lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
#!/bin/bash&lt;br /&gt;
#&lt;br /&gt;
#SBATCH --job-name=&amp;lt;job name&amp;gt;&lt;br /&gt;
#SBATCH --output=&amp;lt;out name&amp;gt;&lt;br /&gt;
#&lt;br /&gt;
#SBATCH --ntasks=&amp;lt;number of tasks&amp;gt;&lt;br /&gt;
#SBATCH --time=&amp;lt;time needed to run task&amp;gt;&lt;br /&gt;
#SBATCH --mem-per-cpu=&amp;lt;memory per cpu needed&amp;gt;&lt;br /&gt;
#SBATCH --mem-per-cpu=&amp;lt;memory per cpu needed&amp;gt;&lt;br /&gt;
#SBATCH --cpus-per-task=&amp;lt;number of cpus&amp;gt;&lt;br /&gt;
&amp;lt;Bash scripting goes here &amp;gt; &lt;br /&gt;
&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This is basically what you need to submit a job to the cluster system we have at hunter. Now lets go through it, shall we?&lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
#SBATCH --job-name=&amp;lt;job_name&amp;gt;&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
This line of code allows you to specify the name of your job, for example if I want my job name to be &amp;quot;pandas_are_awesome&amp;quot; then I would use the following line:&lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
#SBATCH --job-name=pandas_are_awesome&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
The next line of code let you specify the file you want to dump all output (both stdout and stderr) to the terminal into. &#039;&#039;&#039;Remember we are not in interactive mode here, so we wont see any terminal output&#039;&#039;&#039;. So if we want to dump the output to pandas_are_not_awesome.log we would use the following line: &lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
#SBATCH --output=pandas_are_not_awesome.log&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
Hopefully the rest of the #SBATCH lines are self explanatory. &lt;br /&gt;
After we put in what is essentially metadata about our job, we can then write a plain bash script to run our job. So if I wanted to run blast I would use the following code: &lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
#!/bin/bash&lt;br /&gt;
#&lt;br /&gt;
#SBATCH --job-name=blasttest&lt;br /&gt;
#SBATCH --output=blasttest.log&lt;br /&gt;
#&lt;br /&gt;
#SBATCH --ntasks=1&lt;br /&gt;
#SBATCH --time=10:00 #10 minutes &lt;br /&gt;
#SBATCH --mem-per-cpu=100 #megabytes&lt;br /&gt;
#SBATCH --cpus-per-task=4&lt;br /&gt;
&lt;br /&gt;
module load blast/2.2.31 ##load the module &lt;br /&gt;
blastp -query test.fas -db nr -remote -outfmt 6 -out blasttest.out ##run the program &lt;br /&gt;
&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
Notice two things:&lt;br /&gt;
#I needed to load the blast module inorder for my code to run, you need to do this if your program needs to be loaded in order for it to run.&lt;br /&gt;
#You did not have to write &amp;quot;srun&amp;quot; before the command. &lt;br /&gt;
&lt;br /&gt;
Since after the #SBATCH lines, the script becomes a bash scrips, &#039;&#039;&#039;we can also use loops&#039;&#039;&#039;. For example say I had a bunch of sequences in a folder that I wanted to blast; I would use this code: &lt;br /&gt;
&lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
#!/bin/bash&lt;br /&gt;
#&lt;br /&gt;
#SBATCH --job-name=blasttest&lt;br /&gt;
#SBATCH --output=blasttest.log&lt;br /&gt;
#&lt;br /&gt;
#SBATCH --ntasks=1&lt;br /&gt;
#SBATCH --time=10:00 #10 minutes &lt;br /&gt;
#SBATCH --mem-per-cpu=100 #megabytes&lt;br /&gt;
$SBATCH --cpus-per-task=4&lt;br /&gt;
&lt;br /&gt;
module load blast/2.2.31 ##load the module &lt;br /&gt;
for file in ./*.fasta ##go through every file that ends in .fasta&lt;br /&gt;
do &lt;br /&gt;
name=$(basename $file &amp;quot;.fasta&amp;quot; ) ## get the file name &lt;br /&gt;
blastp -query $file -db nr -remote -outfmt 6 -out $name.out ##run blast on the current file &lt;br /&gt;
done&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
Cool beans? Now, after we finish writing and saving the script, how do we run it exactly? &lt;br /&gt;
To do this we use the command on the terminal: &lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
$&amp;gt; sbatch &amp;lt;script.sh&amp;gt; &lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
So if the above script was called &amp;quot;test.sh&amp;quot; to run it I would use the command: &lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
$&amp;gt; sbatch test.sh&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
And it should run without any problems and is now safe to close the connection to the cluster. If we wanted to check on our job there are two possible ways to do it: the not so cool way of looking at our output file we specified in the script. (In this case blasttest.log ) or the cooler ways seen here: &lt;br /&gt;
&lt;br /&gt;
==Monitoring Cluster Jobs==&lt;br /&gt;
To check if your cluster job is running we use the command: &lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
$&amp;gt; squeue&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
Our output would be something like this:&lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
            JOBID PARTITION     NAME     USER ST       TIME  NODES NODELIST(REASON)&lt;br /&gt;
               164      defq     bash     root  R   17:45:29      1 compute003&lt;br /&gt;
               165      defq     bash     root  R   17:45:11      1 compute004&lt;br /&gt;
               168      defq     bash  averdes  R   17:42:19      1 compute006&lt;br /&gt;
               163     himem     bash  jgorson  R 1-17:51:33      1 compute018&lt;br /&gt;
               170     himem     bash  rrahman  blastp 1-17:51:33 1 compute018&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
Notice, we can not only see our own jobs, but also the jobs of the other people on the cluster, so if you are doing something that is incredibly computationally intensive and slowing everyone else down, we will see it and we will bring pitchforks. &lt;br /&gt;
We can also use the command: &lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
$&amp;gt; sstat &lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
To know whats going on with our job, like how much ram or cpu threads it&#039;s using, etc. &lt;br /&gt;
 &lt;br /&gt;
If we want to kill our job, for whatever reason, we would use the command:&lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
$&amp;gt; scancel &amp;lt;jobid&amp;gt;&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
To get our job id, we use squeue, its the first column.&lt;br /&gt;
 &lt;br /&gt;
==Cluster File System and General Information==&lt;br /&gt;
To login to the cluster we ssh into &amp;lt;font color=&amp;quot;red&amp;quot;; size=&amp;quot;24&amp;quot;&amp;gt;146.95.252.94&amp;lt;/font&amp;gt;.&lt;br /&gt;
This can only be done within Hunter, not outside of the firewall. &lt;br /&gt;
&lt;br /&gt;
The cluster file system is separated between 3 areas for storing files and data:&lt;br /&gt;
*&#039;&#039;&#039;Scratch&#039;&#039;&#039;: Scratch space on a file system refers to a shared allocation for &#039;&#039;&#039;all&#039;&#039;&#039; users in the cluster system. Scratch space is reserved for  &amp;lt;font size=&amp;quot;4&amp;quot;&amp;gt;&#039;&#039;&#039;temporary files&#039;&#039;&#039;&amp;lt;/font&amp;gt;. The amount of space allocated here is &#039;&#039;&#039;virtually unlimited&#039;&#039;&#039; however &amp;lt;font size=&amp;quot;4&amp;quot;&amp;gt;&#039;&#039;&#039;files in the scratch directory are purged (deleted) every month&#039;&#039;&#039;&amp;lt;/font&amp;gt;.&lt;br /&gt;
**The directory for the scratch space in Hunters Cluster System is:  &lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
/scratch/ &lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;Personal&#039;&#039;&#039;: You have access to about 10GB of space in your home directory. However since it is your home directory on the cluster on one else can edit or read your files. &amp;lt;font size=&amp;quot;4&amp;quot;&amp;gt;&#039;&#039;&#039;Files here are not purged&#039;&#039;&#039; &amp;lt;/font&amp;gt;&lt;br /&gt;
**The directory for your personal space in Hunters Cluster System is:   &lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
~/&lt;br /&gt;
or&lt;br /&gt;
/home/&amp;lt;username&amp;gt; &lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
*&#039;&#039;&#039;Projects&#039;&#039;&#039;: The project directory is the directory where you should store files related to the work you are doing. Here everyone in your group can view and edit files. Additionally your group has allocations to about 1TB of storage space, collectively. Finally  &amp;lt;font size=&amp;quot;4&amp;quot;&amp;gt;&#039;&#039;&#039;Files here are not purged&#039;&#039;&#039; &amp;lt;/font&amp;gt;&lt;br /&gt;
**The directory for your projects space in Hunters Cluster System is:  &lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
/lustre/projects/&amp;lt;groupname&amp;gt; &lt;br /&gt;
in our case this is &lt;br /&gt;
/lustre/projects/qiulab&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
==Conclusion==&lt;br /&gt;
Hopefully after following this tutorial you have risen from cluster &lt;br /&gt;
Padawan, to cluster Jedi. But complete, your training is not. There&#039;s always something new you can learn about using the cluster system from google or stack overflow. If you find something new or interesting that you think might be useful for other people to use email me at  rayees.rahman40@myhunter.cuny.edu and I can add it to this write up. &lt;br /&gt;
&lt;br /&gt;
This write up is far from complete (we don&#039;t even have a name for our cluster yet, come on!) so new things will be added all the time, so watch this page! &lt;br /&gt;
&lt;br /&gt;
Rayees&lt;br /&gt;
&lt;br /&gt;
=Conda Usage=&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
# to create an environment&lt;br /&gt;
module load anaconda&lt;br /&gt;
conda create --name gbs python=3.10 bwa samtools&lt;br /&gt;
conda info --envs&lt;br /&gt;
&lt;br /&gt;
#  to run&lt;br /&gt;
module load anaconda&lt;br /&gt;
source activate gbs&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
=Pre-written cluster scripts for various software=&lt;br /&gt;
&lt;br /&gt;
Here are examples of scripts written for the cluster that use various software, feel free to use these script to learn how to run particular pieces of bioinformatics software or to create more expansive genomics pipelines. &lt;br /&gt;
Also please feel free to contribute to this page! Send me or Dr. Qiu an email to update this section with new example scripts for running software. &lt;br /&gt;
&lt;br /&gt;
==BWA &amp;amp; Samtools==&lt;br /&gt;
Use array option to loop through directories and files&lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
#!/bin/bash&lt;br /&gt;
#&lt;br /&gt;
#SBATCH --job-name=gbs&lt;br /&gt;
#SBATCH --output=gbs.log&lt;br /&gt;
#SBATCH --array=1-2&lt;br /&gt;
#SBATCH -o output-%A_%a-%J.o&lt;br /&gt;
#SBATCH -n 1&lt;br /&gt;
#&lt;br /&gt;
&lt;br /&gt;
# Reference for array operations: https://portal.supercomputing.wales/index.php/index/slurm/interactive-use-job-arrays/job-arrays/&lt;br /&gt;
# formal doc: https://slurm.schedmd.com/job_array.html&lt;br /&gt;
# %J: job identifier&lt;br /&gt;
# %A: parent job&lt;br /&gt;
# %a: array iteration index&lt;br /&gt;
work_dir=/home/wqiu/project-home/GBS-April-2023&lt;br /&gt;
module load anaconda&lt;br /&gt;
source activate gbs&lt;br /&gt;
echo SLURM_JOB_ID $SLURM_JOB_ID&lt;br /&gt;
echo SLURM_ARRAY_JOB_ID $SLURM_ARRAY_JOB_ID&lt;br /&gt;
echo SLURM_ARRAY_TASK_ID $SLURM_ARRAY_TASK_ID&lt;br /&gt;
&lt;br /&gt;
DIRS=(GBS100 GBS101)&lt;br /&gt;
&lt;br /&gt;
for i in ${DIRS[@]}&lt;br /&gt;
do&lt;br /&gt;
    echo &amp;quot;Processing folder $i&amp;quot;&lt;br /&gt;
    dir=$work_dir/$i&lt;br /&gt;
    bwa mem $work_dir/ref.fas $dir/$i.R1.fq.gz $dir/$i.R2.fq.gz &amp;gt; $dir/$i.sam 2&amp;gt; /dev/null&lt;br /&gt;
    echo &amp;quot;sam file generated: $i.sam&amp;quot;&lt;br /&gt;
    samtools view -bT $work_dir/ref.fas $dir/${i}.sam | samtools sort -o $dir/${i}.bam 2&amp;gt; /dev/null&lt;br /&gt;
    echo &amp;quot;bam file generated: $i.bam&amp;quot;&lt;br /&gt;
done&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Interproscan==&lt;br /&gt;
&lt;br /&gt;
Change the wall clock limit to whatever you feel is appropriate, interproscan runs, as far as I can tell, using only a single core, so it will be pretty slow. Let me know if we can improve the speed of it. &lt;br /&gt;
-Rayes  &lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
#!/bin/bash&lt;br /&gt;
#&lt;br /&gt;
#SBATCH --job-name=interproscan&lt;br /&gt;
#SBATCH --output=interpro.log&lt;br /&gt;
#&lt;br /&gt;
#SBATCH --ntasks=1&lt;br /&gt;
#SBATCH --time=4320:00&lt;br /&gt;
#SBATCH --mem-per-cpu=3000&lt;br /&gt;
#SBATCH --cpus-per-task=4&lt;br /&gt;
&lt;br /&gt;
module load interproscan/5.14.53.0&lt;br /&gt;
interproscan.sh -i cdhit-40 -o cdhit-40-out.txt -t p -goterms -pa -f tsv&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Velvet==&lt;br /&gt;
&lt;br /&gt;
Change the wall clock limit to whatever you feel is appropriate, interproscan runs, as far as I can tell, using only a single core, so it will be pretty slow. Let me know if we can improve the speed of it. &lt;br /&gt;
-Rayes  &lt;br /&gt;
&amp;lt;div class=&amp;quot;toccolours mw-collapsible&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;quot;&amp;gt;&lt;br /&gt;
#!/bin/bash&lt;br /&gt;
#&lt;br /&gt;
#SBATCH --job-name=interproscan&lt;br /&gt;
#SBATCH --output=interpro.log&lt;br /&gt;
#&lt;br /&gt;
#SBATCH --ntasks=1&lt;br /&gt;
#SBATCH --time=4320:00&lt;br /&gt;
#SBATCH --mem-per-cpu=3000&lt;br /&gt;
#SBATCH --cpus-per-task=24&lt;br /&gt;
module load velvet/1.2.10&lt;br /&gt;
velveth ./test2 21 -fastq -long ./N18_S15_L001_R1_001_mat.fq&lt;br /&gt;
velvetg ./test2 -cov_cutoff 4 -min_contig_lgth 100&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=Setting up a conda/mamba environment on the HPC=&lt;br /&gt;
==Create a space for your installation==&lt;br /&gt;
Use the project space since the quota enforced on your home directory is quite small.&lt;br /&gt;
&amp;lt;code&amp;gt;mkdir -p /lustre/projects/qiulab/$USER&amp;lt;/code&amp;gt;&lt;br /&gt;
==Setting up mambaforge==&lt;br /&gt;
Follow the instructions on the [miniforge git repository](https://github.com/conda-forge/miniforge#mambaforge) (partially reproduced below):&lt;br /&gt;
&amp;lt;code&amp;gt;&lt;br /&gt;
curl -L -O &amp;quot;https://github.com/conda-forge/miniforge/releases/latest/download/Mambaforge-$(uname)-$(uname -m).sh&amp;quot;&lt;br /&gt;
bash Mambaforge-$(uname)-$(uname -m).sh&lt;br /&gt;
&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
During the setup wizard, you will be asked to enter the directory in which to install Mambaforge. Enter the directory you created above and add &amp;quot;/mambaforge&amp;quot; to the end. For example, if your username is &amp;quot;user&amp;quot;, you would enter:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;&lt;br /&gt;
/lustre/projects/qiulab/user/mambaforge&lt;br /&gt;
&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Replace &amp;quot;user&amp;quot; above with your own username to use the directory created earlier.&lt;br /&gt;
&lt;br /&gt;
At the end of the installation, you will be asked whether or not you would like to run `conda init`. Answer &amp;quot;yes&amp;quot;. When complete, close and reopen your shell (or disconnect/reconnect to the HPC node).&lt;br /&gt;
&lt;br /&gt;
==Configuring bioconda==&lt;br /&gt;
&lt;br /&gt;
Bioconda(https://bioconda.github.io/) is a conda repository for software used in biology. You&#039;ll need to add this repository for most of the tools you will be using.&lt;br /&gt;
&lt;br /&gt;
Check the website for the latest instructions, but the commands to add the repository (as of Jun, 2023) are below:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;syntaxhighlight lang=bash&amp;gt;&lt;br /&gt;
conda config --add channels defaults&lt;br /&gt;
conda config --add channels bioconda&lt;br /&gt;
conda config --add channels conda-forge&lt;br /&gt;
conda config --set channel_priority strict&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Creating a new environment==&lt;br /&gt;
You can create environments using &amp;lt;code&amp;gt;mamba create&amp;lt;/code&amp;gt; (or &amp;lt;code&amp;gt;conda create&amp;lt;/code&amp;gt;). An environment can be used to install all of the software needed to complete some task or workflow, without interfering with the system installation. You will still be able to access software installed to your operating system&#039;s default PATH, but software in an environment (usually) takes precedent, and is isolated to that environment. In other words, if you install a copy of &amp;lt;code&amp;gt;clustalw&amp;lt;/code&amp;gt; in a conda environment, and you activate that environment, it will (should) mask whatever version of &amp;lt;code&amp;gt;clustalw&amp;lt;/code&amp;gt; you may have installed in your system. Also, you can have more than one version of &amp;lt;code&amp;gt;clustalw&amp;lt;/code&amp;gt; installed simultaneously, each one being installed in its own environment.&lt;br /&gt;
&lt;br /&gt;
===From scratch===&lt;br /&gt;
You can create a new enviroment using &amp;lt;code&amp;gt;mamba create&amp;lt;/code&amp;gt; and supplying a name for the environment. Optionally (but strongly recommended), you can end the command with a list of packages. e.g.,&lt;br /&gt;
&amp;lt;code&amp;gt;&lt;br /&gt;
mamba create aligners -n blast clustalo muscle bwa minimap2 mafft&lt;br /&gt;
&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This will create a new environment called &amp;quot;aligners&amp;quot; with blast, clustalo, muscle, bwa, minimap2, and mafft available for use once activated. To activate an environment, do:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;&lt;br /&gt;
mamba activate aligners&lt;br /&gt;
&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Post-environment creation, you can install packages using the &amp;lt;code&amp;gt;install&amp;lt;/code&amp;gt; subcommand:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;&lt;br /&gt;
mamba install -n aligners diamond&lt;br /&gt;
&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
You can also install specific versions of a software package:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;&lt;br /&gt;
mamba install -n aligners blast=2.10&lt;br /&gt;
&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The above will get conda/mamba to try to downgrade the package, if it is already installed. This may or may not work if the system cannot resolve any dependency issues that may occur with dependent packages.&lt;br /&gt;
&lt;br /&gt;
===From an environment.yml file===&lt;br /&gt;
&lt;br /&gt;
Conda users can share their environments in a fairly convenient fashion by dumping the list of packages:&lt;br /&gt;
&amp;lt;code&amp;gt;&lt;br /&gt;
conda env export -n aligners &amp;gt; aligners_environment.yml&lt;br /&gt;
&amp;lt;/code&amp;gt;&lt;br /&gt;
This will export the list of packages installed in the `aligners`, along with their exact version, into a YAML file called &amp;lt;code&amp;gt;aligners_environment.yml&amp;lt;/code&amp;gt;. You can share this file with others and conda/mamba will attempt to recreate the environment on their machine.&lt;br /&gt;
&lt;br /&gt;
To import a YAML file you recieved, run:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;&lt;br /&gt;
mamba env create -f environment.yml&lt;br /&gt;
&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This will create a new environment with whatever name was specified in the environment.yml file. To choose your own, add &amp;lt;code&amp;gt;-n envname&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
==Using conda/mamba==&lt;br /&gt;
See the &amp;quot;From scratch&amp;quot; section above for a few examples.&lt;br /&gt;
&lt;br /&gt;
Installing Qiu-lab software&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;&lt;br /&gt;
mamba create -n qiulab perl-bio-bpwrapper&lt;br /&gt;
&amp;lt;/code&amp;gt;&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Main_Page&amp;diff=6700</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Main_Page&amp;diff=6700"/>
		<updated>2026-06-22T21:55:13Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{|&lt;br /&gt;
|-&lt;br /&gt;
| &lt;br /&gt;
[[File:Protrait-July-2024-Qiu.jpg| x300 px | thumb |&lt;br /&gt;
&amp;lt;strong&amp;gt;Welcome to Qiu Lab Wiki @ Hunter&amp;lt;/strong&amp;gt;&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;Weigang Qiu, Ph.D., Professor, [https://biology.hunter.cuny.edu Department of Biological Sciences]&amp;lt;br&amp;gt; &lt;br /&gt;
[https://hunter.cuny.edu Hunter College] of [https://www.cuny.edu City University of New York]&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
Adjunct Faculty, Weill Cornell Medical College, [https://weill.cornell.edu/units/systems-and-computational-biomedicine Department of Systems and Computational Biomedicine]&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
Belfer Research Building, Room 402&amp;lt;br&amp;gt;413 East  69th Street, New York, NY 10021&amp;lt;br&amp;gt;Office: 1-212-896-0445&amp;lt;br&amp;gt;Email: wqiu-at-(hunter.cuny.edu) &amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
[https://goo.gl/maps/xv1KmaW3XEnxaY1V7 Directions by Google Map]]]&lt;br /&gt;
|&lt;br /&gt;
[[File:Book-cover.jpg | x300 px | thumb | A textbook for microbial genomics &amp;amp; evolution (Spring 2026):&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;[https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 Read excerpts &amp;amp; Order from the publisher using Author Discount (25%) Code &amp;quot;C2508&amp;quot;]&amp;lt;br&amp;gt;&amp;lt;br&amp;gt; [https://digitalgenomes.net Companion codebook (work in progress &amp;amp;#x1F60A;)]]]&lt;br /&gt;
|&lt;br /&gt;
__TOC__&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Fieldwork Gallery==&lt;br /&gt;
&amp;lt;gallery heights=&amp;quot;200&amp;quot; perrow=&amp;quot;4&amp;quot; style=&amp;quot;text-align:left&amp;quot; widths=&amp;quot;250&amp;quot; mode=&amp;quot;packed&amp;quot;&amp;gt;&lt;br /&gt;
File:Tick trip rockefeller.jpeg|Apr 2026, Rockefeller State Park, NY&lt;br /&gt;
File:Tick-trip-2023.jpg|Oct 2023, Long Island, NY&lt;br /&gt;
File:Tick-trip-2022.jpg|June 2022, MA&lt;br /&gt;
File:Tick-trip-2021.jpg|June 2021, MA&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&amp;lt;youtube width=&amp;quot;400&amp;quot; height=&amp;quot;300&amp;quot;&amp;gt;Jb4ACK-GjM0&amp;lt;/youtube&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Lab publications==&lt;br /&gt;
===Lyme Genomics, Evolution, &amp;amp; Ecology===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;4&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
&lt;br /&gt;
File:Screenshot_2024-08-15_094908.png | link=https://journals.asm.org/doi/10.1128/mbio.01749-24 | Akther et al. 2024. &amp;quot;Natural selection and recombination at host-interacting lipoprotein loci drive genome diversification of Lyme disease and related bacteria&amp;quot; &#039;&#039;&#039;&#039;&#039;mBio&#039;&#039;&#039;&#039;&#039; 0:e01749-24. Press coverage: [https://www.gc.cuny.edu/news/cuny-graduate-center-biologists-map-dna-lyme-disease-bacteria CUNY Graduate Center News Story]; [https://hunter.cuny.edu/news/hunter-researcher-maps-dna-of-lyme-disease-bacteria/ Hunter press]&lt;br /&gt;
&lt;br /&gt;
File:Applsci-13-11587-g004.png | link=https://doi.org/10.3390/app132011587 | Di, Chong, Brian Sulkow, Weigang Qiu, and Shipeng Sun. 2023. &amp;quot;Effects of Micro-Scale Environmental Factors on the Quantity of Questing Black-Legged Ticks in Suburban New York&amp;quot; &#039;&#039;&#039;&#039;&#039;Applied Sciences&#039;&#039;&#039;&#039;&#039; 13, no. 20: 11587.&lt;br /&gt;
&lt;br /&gt;
File:Spectrum.01743-22-f003.gif | link=https://doi.org/10.1128/spectrum.01743-22  | Li, Di, Zeglis, Qiu (2022). “Evolution of the &#039;&#039;vls&#039;&#039; antigenic variability locus of the Lyme Disease pathogen and development of recombinant monoclonal antibodies targeting conserved VlsE epitopes”. &#039;&#039;&#039;&#039;&#039;Microbial Spectrum&#039;&#039;&#039;&#039;&#039; 10 (5):1-15. &lt;br /&gt;
&lt;br /&gt;
File:Fig7-small.png | link= https://pubmed.ncbi.nlm.nih.gov/34413477 | Di*, Akther*, Bezrucenkovas, Ivanova, Sulkow, Wu, Mneimneh, Gomes-Solecki, Qiu (2021). &amp;quot;Maximum antigen diversification in a lyme bacterial population and evolutionary strategies to overcome pathogen diversity&amp;quot;. &#039;&#039;&#039;&#039;&#039;The ISME Journal&#039;&#039;&#039;&#039;&#039;. 16, 447-464. (*co-first authors) [https://communities.springernature.com/posts/jenner-s-dilemma-and-how-to-win-evolutionary-arms-races-against-microbial-pathogens Blog Post]&lt;br /&gt;
&lt;br /&gt;
File:Ira-fig2.png | link=https://doi.org/10.21775/9781913652616 | Schwartz, Margos, Casjens, Qiu, Eggers (2020). &amp;quot;Multipartite Genome of Lyme Disease &#039;&#039;Borrelia&#039;&#039;: Structure, Variation and Prophages&amp;quot;. &#039;&#039;&#039;&#039;&#039;Current Issues in Molecular Biology&#039;&#039;&#039;&#039;&#039;. 42:409-454. &lt;br /&gt;
&lt;br /&gt;
File:Fig6-1-Barbour.png | link=https://doi.org/10.1002/9781118960608.gbm01525 | Barbour &amp;amp; Qiu (2019). &#039;&#039;Borreliella&#039;&#039;. In &#039;&#039;&#039;&#039;&#039;Bergey&#039;s Manual of Systematics of Archaea and Bacteria&#039;&#039;&#039;&#039;&#039;. John Wiley &amp;amp; Sons, Inc., in association with Bergey&#039;s Manual Trust. &lt;br /&gt;
&lt;br /&gt;
File:Zjm9990961420002.jpeg | link=https://doi.org/10.1128/jcm.00940-18 | Di L, Wan Z, Akther S, Ying CX, Larracuente A, Li L, Di C, Nunez R, Cucura DM, Goddard NL, Krampis K, Qiu WG. (2018). Genotyping and Quantifying Lyme Pathogen Strains by Deep Sequencing of the Outer Surface Protein C (&#039;&#039;ospC&#039;&#039;) Locus. &#039;&#039;&#039;Journal of Clinical Microbiology&#039;&#039;&#039;. 56(11):e00940-18. &lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Evolution &amp;amp; Learning Algorithms===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;3&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File:Pathogens-12-00388-g003.png | link=https://doi.org/10.3390/pathogens12030388 | Ely, Koh, Ho, Hassan, Pham, Qiu (2023). Novelty Search Promotes Antigenic Diversity in Microbial Pathogens. &#039;&#039;&#039;&#039;&#039;Pathogens&#039;&#039;&#039;&#039;&#039; 12:388.&lt;br /&gt;
File:Pone.0224552.g005.png | link=https://doi.org/10.1371/journal.pone.0224552 | Attie, Sulkow, Di, Qiu (2019). Genetic codes optimized as a traveling salesman problem. &#039;&#039;&#039;&#039;&#039;PLoS ONE&#039;&#039;&#039;&#039;&#039; 14(10): e0224552.&lt;br /&gt;
File:Pcbi.1005677.g001.png | link=https://doi.org/10.1371/journal.pcbi.1005677 | Yan, Deforet, Boyle, Rahman, Liang, Okegbe, Dietrich, Qiu, Xavier (2017). Bow-tie signaling in c-di-GMP: Machine learning in a simple biochemical network. &#039;&#039;&#039;&#039;&#039;PLoS Comput Biol&#039;&#039;&#039;&#039;&#039; 13(8): e1005677. &lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Informatics Tool Development===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File: 12859 2018 2074 Fig1 HTML.png | link=https://pubmed.ncbi.nlm.nih.gov/29499649 | Hernandez, Bernstein, Pagan, Vargas, McCaig, Ramrattan, Akther, Larracuente, Di, Vieira, Qiu. (2018). BbWrapper: BioPerl-based sequence and tree utilities for rapid prototyping of bioinformatics pipelines. &#039;&#039;&#039;&#039;&#039;BMC Bioinformatics&#039;&#039;&#039;&#039;&#039; 19(1):76.&lt;br /&gt;
File: 12859 2014 Article 6488 Fig3 HTML.png | link=https://pubmed.ncbi.nlm.nih.gov/24994456 | Di, Pagan, Packer, Martin, Akther, Ramrattan, Mongodin, Fraser, Schutzer, Luft, Casjens and Qiu. (2014). BorreliaBase: a phylogeny-centered browser of &#039;&#039;Borrelia&#039;&#039; genomes. &#039;&#039;&#039;&#039;&#039;BMC Bioinformatics&#039;&#039;&#039;&#039;&#039; 15:233.&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*[http://scholar.google.com/citations?hl=en&amp;amp;user=Ds6u39QAAAAJ) Full list by Google Scholar]&lt;br /&gt;
*[http://www.ncbi.nlm.nih.gov/sites/myncbi/weigang.qiu.1/bibliography/42770924/public/ Full list by NCBI Bibliography]&lt;br /&gt;
last update: March 20, 2023&lt;br /&gt;
&lt;br /&gt;
==Lab members and trainees==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
!Year/Period !!Doctoral members &amp;amp; trainees !!Other members &amp;amp; trainees&lt;br /&gt;
|-&lt;br /&gt;
|Current Academic Year&lt;br /&gt;
(Spring &amp;amp; Summer 2026)&lt;br /&gt;
||&lt;br /&gt;
*Edee Amin: CUNY Graduate Center, Bio/EEB doctoral program&lt;br /&gt;
* Dr Yozen Hernandez: System administrator (part-time), Ph.D. from Boston University&lt;br /&gt;
||&lt;br /&gt;
*Tasmina Hassan: Hunter Bio/CS&lt;br /&gt;
*Mathew DiCicco: Hunter Math/CS&lt;br /&gt;
*Tara Doma Lama: Hunter Bio/Bioinformatics&lt;br /&gt;
*Hagar Abuzaid: Hunter Bio&lt;br /&gt;
|-&lt;br /&gt;
|Alumni (Since Fall 2002)&lt;br /&gt;
||&lt;br /&gt;
* Brandon Ely (2026): CUNY Grad Center, Biology/MCD doctoral program&lt;br /&gt;
* Li Li (Lily, 2023): CUNY Grad Center, Biology/EEB doctoral program&lt;br /&gt;
*Dr Lia Di: Ph.D. Postdoctoral Research Associate, from Nanjing Agricultural University &amp;amp; Wisconsin Blood Institute&lt;br /&gt;
*Dr Saymon Akther (2022): CUNY Grad Center, Biology/EEB&lt;br /&gt;
*Dr Rayees Rahman: Hunter Bio/Bioinformatics, Ph.D. from Mt Sinai Medical School&lt;br /&gt;
*Dr Che Martin (2013): CUNY Grad Center, Biology/MCD&lt;br /&gt;
*Dr James Haven (2011): CUNY Grad Center, Biology/MCD&lt;br /&gt;
*Dr Tika Sukarna (2009): CUNY Grad Center, Biology/MCD&lt;br /&gt;
*Dr Juan Coronado (2008): CUNY Grad Center, Biology/MCD (Dr Peter Lipke)&lt;br /&gt;
*Dr William McCaig: CUNY BA, Ph.D. from Stony Brook University&lt;br /&gt;
*Dr Vincent Xue: CUNY CS/Bioinformatics, Ph.D. from MIT&lt;br /&gt;
*Dr Fubin Li: CUNY Grad Center, Biology/MCD (Dr Laurel Eckhardt)&lt;br /&gt;
*Dr Oliver Attie: Postdoctoral Research Associate, Ph.D. from NYU&lt;br /&gt;
||&lt;br /&gt;
(published coauthors)&amp;lt;br&amp;gt; &lt;br /&gt;
*Brian Sulkow: CUNY Grad Center, Mathematics&lt;br /&gt;
* Winston Koh: Hunter Bio/CS&lt;br /&gt;
*Eamen Ho: Hunter Bio/Bioinformatics&lt;br /&gt;
*Ahn Pham: Hunter Bio/Bioinformatics&lt;br /&gt;
*Chris Panlasigui: Hunter Bio/Bioinformatics&lt;br /&gt;
*Amanda Amanda Larracuente: CUNY Grad Center, Biology/MCD&lt;br /&gt;
* Pedro Pagan: Hunter Bio/Bioinformatics&lt;br /&gt;
*Edgaras Bezrucenkovas: Hunter Chem/Bioinformatics&lt;br /&gt;
*Girish Ramrattan: Hunter Bio/Bioinformatics&lt;br /&gt;
* Levy Vargas: Hunter Bio/Bioinformatics&lt;br /&gt;
*Chong Di: Hunter Geography&lt;br /&gt;
*Roy Nunez: Hunter Bio&lt;br /&gt;
*Mei Wu: CUNY City Tech&lt;br /&gt;
*Desiree Pante: Hunter Bio&lt;br /&gt;
*Saimtun Shipa: Hunter Stat (MA)&lt;br /&gt;
*Bing Wu: Hunter Bio/Biotechnology&lt;br /&gt;
* Svidatoslav Kendall (Slav): Hunter Biology&lt;br /&gt;
*Philip Romov: Hunter CS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Web Apps==&lt;br /&gt;
===Apps with Collaborators (or from published papers)===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;4&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File:Screenshot_2024-06-19_205621.png | link=https://cov.genometracker.org/finalelisa-app/ | Borrelia diagnostic antigens (App developed by Liann Aris-Henry; Data from [https://journals.asm.org/doi/10.1128/jcm.01142-19 Arumugam et al (2019)])&lt;br /&gt;
File:Screenshot_2024-06-18_144931.png | link=https://cov.genometracker.org/borrelia-app/ | Borrelia growth transcriptome (App developed by Laziz Asamov; Data from [https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0164165 Arnold et al (2016)])&lt;br /&gt;
File: Silac.png | link=http://borreliabase.org/~wgqiu/Polotskaia_etal_2014/supp-fig-s1/ | Data &amp;amp; [http://borreliabase.org/~wgqiu/Polotskaia_etal_2014/supp-table-s1/ GSEA] associated with Polotskaia_etal_2014 (with Bargonetti Lab @Hunter)&lt;br /&gt;
File:Screenshot_2023-03-24_122016.png | link=http://borreliabase.org/~wgqiu/clickme-khalikuz/temp-Points.html | Phosphoproteome MDM2KD data set (with Bargonetti Lab @Hunter)&lt;br /&gt;
File:Screenshot_2023-03-24_122312.png | link=http://borreliabase.org/~wgqiu/mpai-v3/ | Genes &amp;amp; Pathways Associated with mTP53 (with Bargonetti Lab @Hunter)&lt;br /&gt;
File:Spombe.png | link=Spombe | S. pombe transcriptomes (with Zhong Lab @Hunter)&lt;br /&gt;
File:Screenshot 2023-03-24 122223.png | link=http://borreliabase.org/~wgqiu/gbs-browser-v3/ | Group B Streptococcus (GBS) genome browser (with Wu Lab @Shenzhen)&lt;br /&gt;
File:Screenshot 2023-03-24 122245.png | link=http://borreliabase.org/~wgqiu/oneKGenome/ | 1K genome (with Ogunwobi Lab @Hunter)&lt;br /&gt;
File:Screenshot 2023-03-24 122450.png | link=http://borreliabase.org/~wgqiu/E_faecalis/ | E faecalis genome browser (with Morales Lab @WCMC)&lt;br /&gt;
File:Screenshot 2023-03-24 122145.png | link=http://borreliabase.org/~wgqiu/carmen-proteomics/ | Proteomics GSEA results (with Melendez Lab @Hunter) &lt;br /&gt;
File:Screenshot 2023-03-24 122051.png | link=http://borreliabase.org/~wgqiu/Tcell-pHis-ACS/ | pHis Protein Browser (with Skolnik lab @NYU)&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Qiu Lab Apps===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;3&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File:Screenshot 2023-03-24 122558.png | link=http://borreliabase.org | Lyme genome browser&lt;br /&gt;
File:Screenshot 2023-03-24 133320.png | link=http://borreliabase.org/~wgqiu/pf-trees/ | Lyme pathogen plasmid partitioning gene trees&lt;br /&gt;
File:Screenshot 2023-03-24 133636.png | link=http://borreliabase.org/vls-finder/ | vls Finder in Lyme pathogen genomes&lt;br /&gt;
File:Screenshot 2023-03-24 122359.png | link=http://borreliabase.org/~wgqiu/code-wheel | Codon Wheel&lt;br /&gt;
File:PhyloView.png | link=http://borreliabase.org/~wgqiu/PhyloView | Co-visualization of a tree with an alignment and character matrix&lt;br /&gt;
File:Screenshot 2023-03-24 122521.png | link=http://borreliabase.org/~wgqiu/ospC-sequencing | OspC amplicon sequencing from ticks&lt;br /&gt;
File:Screenshot 2023-03-24 122340.png | link=http://borreliabase.org/~wgqiu/asm-challenge | Genomic epidemiology of a Salmonella outbreak (ASM Challenge)&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Curricular Development &amp;amp; Bioinformatics/QuBi Advising ==&lt;br /&gt;
*QuBi advisors: Weigang Qiu, Ntino Krampis, Rabindra Mandal (Biology); Saad Mneimeih, Lei Xie (CS); Akira Kawamura (Chem); Dana Sylvan (Math &amp;amp; Stats) &lt;br /&gt;
**Permission for non-Biology majors to take BIOL203 &amp;amp; BIOL425, every Spring&lt;br /&gt;
**Collect names,  major, and IDs to send to course coordinator to grant permission. Waive BIOL10200 pre-reqs for taking BIOL203.&lt;br /&gt;
*Curricular resources:&lt;br /&gt;
**[http://biology.hunter.cuny.edu/index.php?option=com_content&amp;amp;view=article&amp;amp;id=66&amp;amp;Itemid=73 Biology courses and pre-reqs]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/BIO1-BA Hunter Biology Major 1 (including the Bioinformatics Option)]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/CHEM2-BA Hunter Chemistry Major 2 (including the Bioinformatics Option)]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/COMPSCI-BA Hunter Computer Science (including Bioinformatics Concentration)]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/MATH-BA Hunter Mathematics (including the Quantitative Biology Concentration)]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/STATS-BA Hunter Statistics (including the Quantitative Biology Concentration)]&lt;br /&gt;
*QuBi advising:&lt;br /&gt;
**Declaration of Bioinformatics concentration: In-person advising to work out the semester-by-semester courses&lt;br /&gt;
**Approve on department spreadsheet (or send email to  &amp;quot;Samantha Sheppard-Lahiji&amp;quot; and &amp;quot;HTR Bio&amp;quot; &amp;lt;biology@hunter.cuny.edu&amp;gt;)&lt;br /&gt;
**Students should take Bioinformatics-specific electives (8 cred; see Hunter Catalog below), &#039;&#039;&#039;not general electives&#039;&#039;&#039;&lt;br /&gt;
***Examples: Anthrop302 (3 cr); Chem333 (3 cr); BIOL47119 &amp;amp; BIOL47120 (3cr); BIOL48002 (2 cr)&lt;br /&gt;
**Students need to take BIOL48002 (2 cr), which counts towards as research credit, to graduate as honors&lt;br /&gt;
*General advising:&lt;br /&gt;
**~40 students every semester. Send out emails to students. Go through student courses by Email or by appointment&lt;br /&gt;
**Recommend new math courses: &#039;&#039;&#039;MATH15200 &amp;amp; STAT21350&#039;&#039;&#039;&lt;br /&gt;
*Hosting QuBi students in lab&lt;br /&gt;
**This is to enhance the informatics and coding skills of our students&lt;br /&gt;
**Students should register and get &#039;&#039;&#039;BIOL48002&#039;&#039;&#039; credits, which counts towards their elective credits &amp;amp; eligibility for honors&lt;br /&gt;
**3-5 students per semester&lt;br /&gt;
*Outside research opportunities&lt;br /&gt;
**MIT Quantitative Workshop (first week of January, in Boston). Coordination with CS (Saad &amp;amp; Susan Epstein) in Fall&lt;br /&gt;
**Simons Foundation/Flatiron Institute Center for Computational Biology (CCB) Internship program. Open House in Spring&lt;br /&gt;
&lt;br /&gt;
==Course/Lecture syllabi==&lt;br /&gt;
*[[NYRaMP-Informatics-2025|NYRaMP Workshop (August 2025, by Brandon Ely)]]&lt;br /&gt;
*[[Computational Genomics (KIZ, Fall 2024)]]&lt;br /&gt;
*[[NYRaMP-Informatics-2024|NYRaMP Workshop (August 2024)]]&lt;br /&gt;
*BIOL47120 BioMedical Genomics (Spring 2024). Tutorials: [https://borreliabase.org/~wgqiu/tutorial-markdown.html R Markdown] [https://borreliabase.org/~wgqiu/cluster-analysis.html Cluster analysis] [https://borreliabase.org/~wgqiu/scRNA-analysis.html single-cell RNA-seq]&lt;br /&gt;
*BIOL425 Computational Molecular Biology (Spring, 2023). [https://github.com/weigangq/CSB-BIOL425/tree/master/lecture-materials Lecture material on github]&lt;br /&gt;
*BIOL714 Cell Biology: [http://borreliabase.org/~wgqiu/r-demo-2024.html R Demo (Spring 2024)] [http://borreliabase.org/~wgqiu/r-demo-2023.html R Demo (Spring 2023)]&lt;br /&gt;
*QuBi module: [[QuBi/module/bio203-lab12—2022|BIOL20300 Molecular Genetics, Lab 12 (2023)]]&lt;br /&gt;
*QuBi module: [[QuBi/modules/biol203-geno-pheno-association-2022|BIOL20300 Molecular Genetics, Lab 13 (2022)]]&lt;br /&gt;
*QuBi module: [[QuBi/modules/biol303|BIOL30300 Cell Biology, Bioinformatics Lab (transcriptome analysis)]]&lt;br /&gt;
*[[BigData 2020|Big Data (Summer, 2020)]]&lt;br /&gt;
*[[BioMed-R-2020|BIOL47120 Biomedical Genomics II (Spring, 2020)]] &lt;br /&gt;
**[http://borreliabase.org/~wgqiu/tutorial-markdown.html Tutorial: R Markdown (Spring 2024)]&lt;br /&gt;
**[http://borreliabase.org/~wgqiu/cluster-analysis.html Tutorial: Cluster analysis (Spring 2024)]&lt;br /&gt;
**[http://borreliabase.org/~wgqiu/scRNA-analysis.html Tutorial: single-cel transcriptome analysis (Spring 2024)]&lt;br /&gt;
*[[Biol425 2020|BIOL425 Computational Molecular Biology (Spring, 2020)]]&lt;br /&gt;
*[[Biol375 2019|BIOL37500, Molecular Evolution (Fall, 2019)]]&lt;br /&gt;
*[[Southwest-University|Southwest University R course (Summer, 2019)]]&lt;br /&gt;
*[[Biol20N02 2017|Analysis of Biological Data (Spring, 2017)]]&lt;br /&gt;
*[[Bioinformatics_Workshop_2014|Bioinformatics Workshop (Summer, 2014)]]&lt;br /&gt;
&lt;br /&gt;
==SARS-CoV-2 genome evolution==&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;3&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File:Cov-fig1.jpg | Akther, Bezrucenkovas, Sulkow, Panlasigui, Qiu, Di (April, 2020). &amp;quot;CoV Genome Tracker: tracing genomic footprints of Covid-19 pandemic&amp;quot;. &#039;&#039;&#039;&#039;&#039;[https://www.biorxiv.org/content/biorxiv/early/2020/04/14/2020.04.10.036343.full.pdf BioRxiv]&#039;&#039;&#039;&#039;&#039;; [https://cov.genometracker.org/ Web app: SARS-CoV-2 Genome Tracker]; Github: https://github.com/weigangq/cov-browser&lt;br /&gt;
File:Rec-fig2.png | Akther, Li, Martin, Di, Sulkow, Pante, Bezrucenlovas, Luft, Qiu (May, 2020). &amp;quot;Origin, recombination, and missed opprotunities:  a genomic perspective of the first 100 days of COVID-19 pandemic&amp;quot;. (Unpublished).&lt;br /&gt;
File:Cov-fig3-trace.png | Akther, Bezrucenlovas, Li, Sulkow, Di, Pante, Martin, Luft, Qiu (Sep, 2021). &amp;quot;Following the Trail of One Million Genomes: Footprints of SARS-CoV-2 Adaptation to Humans&amp;quot;. &#039;&#039;&#039;&#039;&#039;[https://www.biorxiv.org/content/biorxiv/early/2021/05/10/2021.05.07.443114.full.pdf BioRxiv link]&#039;&#039;&#039;&#039;&#039;: . Github:  https://github.com/weigangq/cov-db&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Lab Resources &amp;amp; Protocols==&lt;br /&gt;
*OneRF usage (Unofficial; login needed): https://sites.google.com/view/onerf/home&lt;br /&gt;
*[https://runestone.academy/ns/books/published/thinkcspy/index.html How to think like a computer scientist: An interactive Python programming book]&lt;br /&gt;
* Nanopore sequencing protocols&lt;br /&gt;
**DNA barcoding: https://nanopore4edu.org/latest/annotated_experiments/dna_barcoding/&lt;br /&gt;
**Yeast genomes: https://nanoporetech.com/document/extraction-method/yeast-dna&lt;br /&gt;
*[[Monte Carlo Club]]&lt;br /&gt;
*[[NY-RaMP Mentoring]]&lt;br /&gt;
*Borreliella genome sequencing consortium: Weekly meetings (Tu @11): Since Jan 2023&lt;br /&gt;
*Borreliella diagnostic antigens (Fall 2023-Fall 2027):&lt;br /&gt;
**Zoom call (Jan 23, 2024)&lt;br /&gt;
** Next meeting: March 23, 2024&lt;br /&gt;
*Qiu lab network [[First Time Guide|first-time user guide]]&lt;br /&gt;
*Qiu lab Github repositories: https://github.com/weigangq/?tab=repositories&lt;br /&gt;
*[[Mini-Tutorals|Mini-Protocols]] (frequently used computer codes and pipelines)&lt;br /&gt;
*Python tutorial: https://wiki.genometracker.org/~weigang/Intro_to_Python.html&lt;br /&gt;
* [[Tick protocol|ick handling protocols]]&lt;br /&gt;
*[[A Primer on the Cluster System at Hunter|Hunter HPC Usage]]&lt;br /&gt;
*[https://r4ds.hadley.nz/ R for Data Science (2e)], (2024) by Wickham, Grolemund &amp;amp; Çetinkaya-Rundel ([https://bookdown.org/ Bookdown version])&lt;br /&gt;
*Borrelia Genome Consortium:  https://www.ncbi.nlm.nih.gov/bioproject/PRJNA431102/&lt;br /&gt;
*Canadian Bbsl genome assemblies: https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA1130942&lt;br /&gt;
* Nanopore sequencing resources:&lt;br /&gt;
**eBook: https://store.nanoporetech.com/us/minion.html&lt;br /&gt;
&lt;br /&gt;
==Wiki Help==&lt;br /&gt;
*[https://www.mediawiki.org/wiki/Special:MyLanguage/Manual:Configuration_settings Configuration settings list]&lt;br /&gt;
*[https://www.mediawiki.org/wiki/Special:MyLanguage/Manual:FAQ MediaWiki FAQ]&lt;br /&gt;
*[https://lists.wikimedia.org/postorius/lists/mediawiki-announce.lists.wikimedia.org/ MediaWiki release mailing list]&lt;br /&gt;
*[https://www.mediawiki.org/wiki/Special:MyLanguage/Localisation#Translation_resources Localise MediaWiki for your language]&lt;br /&gt;
*[https://www.mediawiki.org/wiki/Special:MyLanguage/Manual:Combating_spam Learn how to combat spam on your wiki]&lt;br /&gt;
*Consult the [[mediawikiwiki:Special:MyLanguage/Help:Contents|User&#039;s Guide]] for information on using the wiki software.&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
	<entry>
		<id>https://wiki.genometracker.org/index.php?title=Main_Page&amp;diff=6699</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://wiki.genometracker.org/index.php?title=Main_Page&amp;diff=6699"/>
		<updated>2026-06-22T21:54:18Z</updated>

		<summary type="html">&lt;p&gt;Wikiuser: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{|&lt;br /&gt;
|-&lt;br /&gt;
| &lt;br /&gt;
[[File:Protrait-July-2024-Qiu.jpg| x300 px | thumb |&lt;br /&gt;
&amp;lt;strong&amp;gt;Welcome to Qiu Lab Wiki @ Hunter&amp;lt;/strong&amp;gt;&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;Weigang Qiu, Ph.D., Professor, [https://biology.hunter.cuny.edu Department of Biological Sciences]&amp;lt;br&amp;gt; &lt;br /&gt;
[https://hunter.cuny.edu Hunter College] of [https://www.cuny.edu City University of New York]&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
Adjunct Faculty, Weill Cornell Medical College, [https://weill.cornell.edu/units/systems-and-computational-biomedicine Department of Systems and Computational Biomedicine]&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
Belfer Research Building, Room 402&amp;lt;br&amp;gt;413 East  69th Street, New York, NY 10021&amp;lt;br&amp;gt;Office: 1-212-896-0445&amp;lt;br&amp;gt;Email: wqiu-at-(hunter.cuny.edu) &amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
[https://goo.gl/maps/xv1KmaW3XEnxaY1V7 Directions by Google Map]]]&lt;br /&gt;
|&lt;br /&gt;
[[File:Book-cover.jpg | x300 px | thumb | A textbook for microbial genomics &amp;amp; evolution (Spring 2026):&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;[https://www.wiley.com/en-us/shop/general-introductory-life-sciences/digital-genomes-monte-carlo-simulations-of-microbes-and-evolution-p-9781394314621 Read excerpts &amp;amp; Order from the publisher using Author Discount (25%) Code &amp;quot;C2508&amp;quot;]&amp;lt;br&amp;gt;&amp;lt;br&amp;gt; [https://digitalgenomes.net Companion codebook (work in progress )]]]&lt;br /&gt;
|&lt;br /&gt;
__TOC__&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Fieldwork Gallery==&lt;br /&gt;
&amp;lt;gallery heights=&amp;quot;200&amp;quot; perrow=&amp;quot;4&amp;quot; style=&amp;quot;text-align:left&amp;quot; widths=&amp;quot;250&amp;quot; mode=&amp;quot;packed&amp;quot;&amp;gt;&lt;br /&gt;
File:Tick trip rockefeller.jpeg|Apr 2026, Rockefeller State Park, NY&lt;br /&gt;
File:Tick-trip-2023.jpg|Oct 2023, Long Island, NY&lt;br /&gt;
File:Tick-trip-2022.jpg|June 2022, MA&lt;br /&gt;
File:Tick-trip-2021.jpg|June 2021, MA&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&amp;lt;youtube width=&amp;quot;400&amp;quot; height=&amp;quot;300&amp;quot;&amp;gt;Jb4ACK-GjM0&amp;lt;/youtube&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Lab publications==&lt;br /&gt;
===Lyme Genomics, Evolution, &amp;amp; Ecology===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;4&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
&lt;br /&gt;
File:Screenshot_2024-08-15_094908.png | link=https://journals.asm.org/doi/10.1128/mbio.01749-24 | Akther et al. 2024. &amp;quot;Natural selection and recombination at host-interacting lipoprotein loci drive genome diversification of Lyme disease and related bacteria&amp;quot; &#039;&#039;&#039;&#039;&#039;mBio&#039;&#039;&#039;&#039;&#039; 0:e01749-24. Press coverage: [https://www.gc.cuny.edu/news/cuny-graduate-center-biologists-map-dna-lyme-disease-bacteria CUNY Graduate Center News Story]; [https://hunter.cuny.edu/news/hunter-researcher-maps-dna-of-lyme-disease-bacteria/ Hunter press]&lt;br /&gt;
&lt;br /&gt;
File:Applsci-13-11587-g004.png | link=https://doi.org/10.3390/app132011587 | Di, Chong, Brian Sulkow, Weigang Qiu, and Shipeng Sun. 2023. &amp;quot;Effects of Micro-Scale Environmental Factors on the Quantity of Questing Black-Legged Ticks in Suburban New York&amp;quot; &#039;&#039;&#039;&#039;&#039;Applied Sciences&#039;&#039;&#039;&#039;&#039; 13, no. 20: 11587.&lt;br /&gt;
&lt;br /&gt;
File:Spectrum.01743-22-f003.gif | link=https://doi.org/10.1128/spectrum.01743-22  | Li, Di, Zeglis, Qiu (2022). “Evolution of the &#039;&#039;vls&#039;&#039; antigenic variability locus of the Lyme Disease pathogen and development of recombinant monoclonal antibodies targeting conserved VlsE epitopes”. &#039;&#039;&#039;&#039;&#039;Microbial Spectrum&#039;&#039;&#039;&#039;&#039; 10 (5):1-15. &lt;br /&gt;
&lt;br /&gt;
File:Fig7-small.png | link= https://pubmed.ncbi.nlm.nih.gov/34413477 | Di*, Akther*, Bezrucenkovas, Ivanova, Sulkow, Wu, Mneimneh, Gomes-Solecki, Qiu (2021). &amp;quot;Maximum antigen diversification in a lyme bacterial population and evolutionary strategies to overcome pathogen diversity&amp;quot;. &#039;&#039;&#039;&#039;&#039;The ISME Journal&#039;&#039;&#039;&#039;&#039;. 16, 447-464. (*co-first authors) [https://communities.springernature.com/posts/jenner-s-dilemma-and-how-to-win-evolutionary-arms-races-against-microbial-pathogens Blog Post]&lt;br /&gt;
&lt;br /&gt;
File:Ira-fig2.png | link=https://doi.org/10.21775/9781913652616 | Schwartz, Margos, Casjens, Qiu, Eggers (2020). &amp;quot;Multipartite Genome of Lyme Disease &#039;&#039;Borrelia&#039;&#039;: Structure, Variation and Prophages&amp;quot;. &#039;&#039;&#039;&#039;&#039;Current Issues in Molecular Biology&#039;&#039;&#039;&#039;&#039;. 42:409-454. &lt;br /&gt;
&lt;br /&gt;
File:Fig6-1-Barbour.png | link=https://doi.org/10.1002/9781118960608.gbm01525 | Barbour &amp;amp; Qiu (2019). &#039;&#039;Borreliella&#039;&#039;. In &#039;&#039;&#039;&#039;&#039;Bergey&#039;s Manual of Systematics of Archaea and Bacteria&#039;&#039;&#039;&#039;&#039;. John Wiley &amp;amp; Sons, Inc., in association with Bergey&#039;s Manual Trust. &lt;br /&gt;
&lt;br /&gt;
File:Zjm9990961420002.jpeg | link=https://doi.org/10.1128/jcm.00940-18 | Di L, Wan Z, Akther S, Ying CX, Larracuente A, Li L, Di C, Nunez R, Cucura DM, Goddard NL, Krampis K, Qiu WG. (2018). Genotyping and Quantifying Lyme Pathogen Strains by Deep Sequencing of the Outer Surface Protein C (&#039;&#039;ospC&#039;&#039;) Locus. &#039;&#039;&#039;Journal of Clinical Microbiology&#039;&#039;&#039;. 56(11):e00940-18. &lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Evolution &amp;amp; Learning Algorithms===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;3&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File:Pathogens-12-00388-g003.png | link=https://doi.org/10.3390/pathogens12030388 | Ely, Koh, Ho, Hassan, Pham, Qiu (2023). Novelty Search Promotes Antigenic Diversity in Microbial Pathogens. &#039;&#039;&#039;&#039;&#039;Pathogens&#039;&#039;&#039;&#039;&#039; 12:388.&lt;br /&gt;
File:Pone.0224552.g005.png | link=https://doi.org/10.1371/journal.pone.0224552 | Attie, Sulkow, Di, Qiu (2019). Genetic codes optimized as a traveling salesman problem. &#039;&#039;&#039;&#039;&#039;PLoS ONE&#039;&#039;&#039;&#039;&#039; 14(10): e0224552.&lt;br /&gt;
File:Pcbi.1005677.g001.png | link=https://doi.org/10.1371/journal.pcbi.1005677 | Yan, Deforet, Boyle, Rahman, Liang, Okegbe, Dietrich, Qiu, Xavier (2017). Bow-tie signaling in c-di-GMP: Machine learning in a simple biochemical network. &#039;&#039;&#039;&#039;&#039;PLoS Comput Biol&#039;&#039;&#039;&#039;&#039; 13(8): e1005677. &lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Informatics Tool Development===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File: 12859 2018 2074 Fig1 HTML.png | link=https://pubmed.ncbi.nlm.nih.gov/29499649 | Hernandez, Bernstein, Pagan, Vargas, McCaig, Ramrattan, Akther, Larracuente, Di, Vieira, Qiu. (2018). BbWrapper: BioPerl-based sequence and tree utilities for rapid prototyping of bioinformatics pipelines. &#039;&#039;&#039;&#039;&#039;BMC Bioinformatics&#039;&#039;&#039;&#039;&#039; 19(1):76.&lt;br /&gt;
File: 12859 2014 Article 6488 Fig3 HTML.png | link=https://pubmed.ncbi.nlm.nih.gov/24994456 | Di, Pagan, Packer, Martin, Akther, Ramrattan, Mongodin, Fraser, Schutzer, Luft, Casjens and Qiu. (2014). BorreliaBase: a phylogeny-centered browser of &#039;&#039;Borrelia&#039;&#039; genomes. &#039;&#039;&#039;&#039;&#039;BMC Bioinformatics&#039;&#039;&#039;&#039;&#039; 15:233.&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
*[http://scholar.google.com/citations?hl=en&amp;amp;user=Ds6u39QAAAAJ) Full list by Google Scholar]&lt;br /&gt;
*[http://www.ncbi.nlm.nih.gov/sites/myncbi/weigang.qiu.1/bibliography/42770924/public/ Full list by NCBI Bibliography]&lt;br /&gt;
last update: March 20, 2023&lt;br /&gt;
&lt;br /&gt;
==Lab members and trainees==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
!Year/Period !!Doctoral members &amp;amp; trainees !!Other members &amp;amp; trainees&lt;br /&gt;
|-&lt;br /&gt;
|Current Academic Year&lt;br /&gt;
(Spring &amp;amp; Summer 2026)&lt;br /&gt;
||&lt;br /&gt;
*Edee Amin: CUNY Graduate Center, Bio/EEB doctoral program&lt;br /&gt;
* Dr Yozen Hernandez: System administrator (part-time), Ph.D. from Boston University&lt;br /&gt;
||&lt;br /&gt;
*Tasmina Hassan: Hunter Bio/CS&lt;br /&gt;
*Mathew DiCicco: Hunter Math/CS&lt;br /&gt;
*Tara Doma Lama: Hunter Bio/Bioinformatics&lt;br /&gt;
*Hagar Abuzaid: Hunter Bio&lt;br /&gt;
|-&lt;br /&gt;
|Alumni (Since Fall 2002)&lt;br /&gt;
||&lt;br /&gt;
* Brandon Ely (2026): CUNY Grad Center, Biology/MCD doctoral program&lt;br /&gt;
* Li Li (Lily, 2023): CUNY Grad Center, Biology/EEB doctoral program&lt;br /&gt;
*Dr Lia Di: Ph.D. Postdoctoral Research Associate, from Nanjing Agricultural University &amp;amp; Wisconsin Blood Institute&lt;br /&gt;
*Dr Saymon Akther (2022): CUNY Grad Center, Biology/EEB&lt;br /&gt;
*Dr Rayees Rahman: Hunter Bio/Bioinformatics, Ph.D. from Mt Sinai Medical School&lt;br /&gt;
*Dr Che Martin (2013): CUNY Grad Center, Biology/MCD&lt;br /&gt;
*Dr James Haven (2011): CUNY Grad Center, Biology/MCD&lt;br /&gt;
*Dr Tika Sukarna (2009): CUNY Grad Center, Biology/MCD&lt;br /&gt;
*Dr Juan Coronado (2008): CUNY Grad Center, Biology/MCD (Dr Peter Lipke)&lt;br /&gt;
*Dr William McCaig: CUNY BA, Ph.D. from Stony Brook University&lt;br /&gt;
*Dr Vincent Xue: CUNY CS/Bioinformatics, Ph.D. from MIT&lt;br /&gt;
*Dr Fubin Li: CUNY Grad Center, Biology/MCD (Dr Laurel Eckhardt)&lt;br /&gt;
*Dr Oliver Attie: Postdoctoral Research Associate, Ph.D. from NYU&lt;br /&gt;
||&lt;br /&gt;
(published coauthors)&amp;lt;br&amp;gt; &lt;br /&gt;
*Brian Sulkow: CUNY Grad Center, Mathematics&lt;br /&gt;
* Winston Koh: Hunter Bio/CS&lt;br /&gt;
*Eamen Ho: Hunter Bio/Bioinformatics&lt;br /&gt;
*Ahn Pham: Hunter Bio/Bioinformatics&lt;br /&gt;
*Chris Panlasigui: Hunter Bio/Bioinformatics&lt;br /&gt;
*Amanda Amanda Larracuente: CUNY Grad Center, Biology/MCD&lt;br /&gt;
* Pedro Pagan: Hunter Bio/Bioinformatics&lt;br /&gt;
*Edgaras Bezrucenkovas: Hunter Chem/Bioinformatics&lt;br /&gt;
*Girish Ramrattan: Hunter Bio/Bioinformatics&lt;br /&gt;
* Levy Vargas: Hunter Bio/Bioinformatics&lt;br /&gt;
*Chong Di: Hunter Geography&lt;br /&gt;
*Roy Nunez: Hunter Bio&lt;br /&gt;
*Mei Wu: CUNY City Tech&lt;br /&gt;
*Desiree Pante: Hunter Bio&lt;br /&gt;
*Saimtun Shipa: Hunter Stat (MA)&lt;br /&gt;
*Bing Wu: Hunter Bio/Biotechnology&lt;br /&gt;
* Svidatoslav Kendall (Slav): Hunter Biology&lt;br /&gt;
*Philip Romov: Hunter CS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==Web Apps==&lt;br /&gt;
===Apps with Collaborators (or from published papers)===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;4&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File:Screenshot_2024-06-19_205621.png | link=https://cov.genometracker.org/finalelisa-app/ | Borrelia diagnostic antigens (App developed by Liann Aris-Henry; Data from [https://journals.asm.org/doi/10.1128/jcm.01142-19 Arumugam et al (2019)])&lt;br /&gt;
File:Screenshot_2024-06-18_144931.png | link=https://cov.genometracker.org/borrelia-app/ | Borrelia growth transcriptome (App developed by Laziz Asamov; Data from [https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0164165 Arnold et al (2016)])&lt;br /&gt;
File: Silac.png | link=http://borreliabase.org/~wgqiu/Polotskaia_etal_2014/supp-fig-s1/ | Data &amp;amp; [http://borreliabase.org/~wgqiu/Polotskaia_etal_2014/supp-table-s1/ GSEA] associated with Polotskaia_etal_2014 (with Bargonetti Lab @Hunter)&lt;br /&gt;
File:Screenshot_2023-03-24_122016.png | link=http://borreliabase.org/~wgqiu/clickme-khalikuz/temp-Points.html | Phosphoproteome MDM2KD data set (with Bargonetti Lab @Hunter)&lt;br /&gt;
File:Screenshot_2023-03-24_122312.png | link=http://borreliabase.org/~wgqiu/mpai-v3/ | Genes &amp;amp; Pathways Associated with mTP53 (with Bargonetti Lab @Hunter)&lt;br /&gt;
File:Spombe.png | link=Spombe | S. pombe transcriptomes (with Zhong Lab @Hunter)&lt;br /&gt;
File:Screenshot 2023-03-24 122223.png | link=http://borreliabase.org/~wgqiu/gbs-browser-v3/ | Group B Streptococcus (GBS) genome browser (with Wu Lab @Shenzhen)&lt;br /&gt;
File:Screenshot 2023-03-24 122245.png | link=http://borreliabase.org/~wgqiu/oneKGenome/ | 1K genome (with Ogunwobi Lab @Hunter)&lt;br /&gt;
File:Screenshot 2023-03-24 122450.png | link=http://borreliabase.org/~wgqiu/E_faecalis/ | E faecalis genome browser (with Morales Lab @WCMC)&lt;br /&gt;
File:Screenshot 2023-03-24 122145.png | link=http://borreliabase.org/~wgqiu/carmen-proteomics/ | Proteomics GSEA results (with Melendez Lab @Hunter) &lt;br /&gt;
File:Screenshot 2023-03-24 122051.png | link=http://borreliabase.org/~wgqiu/Tcell-pHis-ACS/ | pHis Protein Browser (with Skolnik lab @NYU)&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Qiu Lab Apps===&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;3&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File:Screenshot 2023-03-24 122558.png | link=http://borreliabase.org | Lyme genome browser&lt;br /&gt;
File:Screenshot 2023-03-24 133320.png | link=http://borreliabase.org/~wgqiu/pf-trees/ | Lyme pathogen plasmid partitioning gene trees&lt;br /&gt;
File:Screenshot 2023-03-24 133636.png | link=http://borreliabase.org/vls-finder/ | vls Finder in Lyme pathogen genomes&lt;br /&gt;
File:Screenshot 2023-03-24 122359.png | link=http://borreliabase.org/~wgqiu/code-wheel | Codon Wheel&lt;br /&gt;
File:PhyloView.png | link=http://borreliabase.org/~wgqiu/PhyloView | Co-visualization of a tree with an alignment and character matrix&lt;br /&gt;
File:Screenshot 2023-03-24 122521.png | link=http://borreliabase.org/~wgqiu/ospC-sequencing | OspC amplicon sequencing from ticks&lt;br /&gt;
File:Screenshot 2023-03-24 122340.png | link=http://borreliabase.org/~wgqiu/asm-challenge | Genomic epidemiology of a Salmonella outbreak (ASM Challenge)&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Curricular Development &amp;amp; Bioinformatics/QuBi Advising ==&lt;br /&gt;
*QuBi advisors: Weigang Qiu, Ntino Krampis, Rabindra Mandal (Biology); Saad Mneimeih, Lei Xie (CS); Akira Kawamura (Chem); Dana Sylvan (Math &amp;amp; Stats) &lt;br /&gt;
**Permission for non-Biology majors to take BIOL203 &amp;amp; BIOL425, every Spring&lt;br /&gt;
**Collect names,  major, and IDs to send to course coordinator to grant permission. Waive BIOL10200 pre-reqs for taking BIOL203.&lt;br /&gt;
*Curricular resources:&lt;br /&gt;
**[http://biology.hunter.cuny.edu/index.php?option=com_content&amp;amp;view=article&amp;amp;id=66&amp;amp;Itemid=73 Biology courses and pre-reqs]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/BIO1-BA Hunter Biology Major 1 (including the Bioinformatics Option)]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/CHEM2-BA Hunter Chemistry Major 2 (including the Bioinformatics Option)]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/COMPSCI-BA Hunter Computer Science (including Bioinformatics Concentration)]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/MATH-BA Hunter Mathematics (including the Quantitative Biology Concentration)]&lt;br /&gt;
**[https://hunter-undergraduate.catalog.cuny.edu/programs/STATS-BA Hunter Statistics (including the Quantitative Biology Concentration)]&lt;br /&gt;
*QuBi advising:&lt;br /&gt;
**Declaration of Bioinformatics concentration: In-person advising to work out the semester-by-semester courses&lt;br /&gt;
**Approve on department spreadsheet (or send email to  &amp;quot;Samantha Sheppard-Lahiji&amp;quot; and &amp;quot;HTR Bio&amp;quot; &amp;lt;biology@hunter.cuny.edu&amp;gt;)&lt;br /&gt;
**Students should take Bioinformatics-specific electives (8 cred; see Hunter Catalog below), &#039;&#039;&#039;not general electives&#039;&#039;&#039;&lt;br /&gt;
***Examples: Anthrop302 (3 cr); Chem333 (3 cr); BIOL47119 &amp;amp; BIOL47120 (3cr); BIOL48002 (2 cr)&lt;br /&gt;
**Students need to take BIOL48002 (2 cr), which counts towards as research credit, to graduate as honors&lt;br /&gt;
*General advising:&lt;br /&gt;
**~40 students every semester. Send out emails to students. Go through student courses by Email or by appointment&lt;br /&gt;
**Recommend new math courses: &#039;&#039;&#039;MATH15200 &amp;amp; STAT21350&#039;&#039;&#039;&lt;br /&gt;
*Hosting QuBi students in lab&lt;br /&gt;
**This is to enhance the informatics and coding skills of our students&lt;br /&gt;
**Students should register and get &#039;&#039;&#039;BIOL48002&#039;&#039;&#039; credits, which counts towards their elective credits &amp;amp; eligibility for honors&lt;br /&gt;
**3-5 students per semester&lt;br /&gt;
*Outside research opportunities&lt;br /&gt;
**MIT Quantitative Workshop (first week of January, in Boston). Coordination with CS (Saad &amp;amp; Susan Epstein) in Fall&lt;br /&gt;
**Simons Foundation/Flatiron Institute Center for Computational Biology (CCB) Internship program. Open House in Spring&lt;br /&gt;
&lt;br /&gt;
==Course/Lecture syllabi==&lt;br /&gt;
*[[NYRaMP-Informatics-2025|NYRaMP Workshop (August 2025, by Brandon Ely)]]&lt;br /&gt;
*[[Computational Genomics (KIZ, Fall 2024)]]&lt;br /&gt;
*[[NYRaMP-Informatics-2024|NYRaMP Workshop (August 2024)]]&lt;br /&gt;
*BIOL47120 BioMedical Genomics (Spring 2024). Tutorials: [https://borreliabase.org/~wgqiu/tutorial-markdown.html R Markdown] [https://borreliabase.org/~wgqiu/cluster-analysis.html Cluster analysis] [https://borreliabase.org/~wgqiu/scRNA-analysis.html single-cell RNA-seq]&lt;br /&gt;
*BIOL425 Computational Molecular Biology (Spring, 2023). [https://github.com/weigangq/CSB-BIOL425/tree/master/lecture-materials Lecture material on github]&lt;br /&gt;
*BIOL714 Cell Biology: [http://borreliabase.org/~wgqiu/r-demo-2024.html R Demo (Spring 2024)] [http://borreliabase.org/~wgqiu/r-demo-2023.html R Demo (Spring 2023)]&lt;br /&gt;
*QuBi module: [[QuBi/module/bio203-lab12—2022|BIOL20300 Molecular Genetics, Lab 12 (2023)]]&lt;br /&gt;
*QuBi module: [[QuBi/modules/biol203-geno-pheno-association-2022|BIOL20300 Molecular Genetics, Lab 13 (2022)]]&lt;br /&gt;
*QuBi module: [[QuBi/modules/biol303|BIOL30300 Cell Biology, Bioinformatics Lab (transcriptome analysis)]]&lt;br /&gt;
*[[BigData 2020|Big Data (Summer, 2020)]]&lt;br /&gt;
*[[BioMed-R-2020|BIOL47120 Biomedical Genomics II (Spring, 2020)]] &lt;br /&gt;
**[http://borreliabase.org/~wgqiu/tutorial-markdown.html Tutorial: R Markdown (Spring 2024)]&lt;br /&gt;
**[http://borreliabase.org/~wgqiu/cluster-analysis.html Tutorial: Cluster analysis (Spring 2024)]&lt;br /&gt;
**[http://borreliabase.org/~wgqiu/scRNA-analysis.html Tutorial: single-cel transcriptome analysis (Spring 2024)]&lt;br /&gt;
*[[Biol425 2020|BIOL425 Computational Molecular Biology (Spring, 2020)]]&lt;br /&gt;
*[[Biol375 2019|BIOL37500, Molecular Evolution (Fall, 2019)]]&lt;br /&gt;
*[[Southwest-University|Southwest University R course (Summer, 2019)]]&lt;br /&gt;
*[[Biol20N02 2017|Analysis of Biological Data (Spring, 2017)]]&lt;br /&gt;
*[[Bioinformatics_Workshop_2014|Bioinformatics Workshop (Summer, 2014)]]&lt;br /&gt;
&lt;br /&gt;
==SARS-CoV-2 genome evolution==&lt;br /&gt;
&amp;lt;gallery mode=&amp;quot;packed&amp;quot; heights=&amp;quot;200px&amp;quot; perrow=&amp;quot;3&amp;quot; style=&amp;quot;text-align:left&amp;quot;&amp;gt;&lt;br /&gt;
File:Cov-fig1.jpg | Akther, Bezrucenkovas, Sulkow, Panlasigui, Qiu, Di (April, 2020). &amp;quot;CoV Genome Tracker: tracing genomic footprints of Covid-19 pandemic&amp;quot;. &#039;&#039;&#039;&#039;&#039;[https://www.biorxiv.org/content/biorxiv/early/2020/04/14/2020.04.10.036343.full.pdf BioRxiv]&#039;&#039;&#039;&#039;&#039;; [https://cov.genometracker.org/ Web app: SARS-CoV-2 Genome Tracker]; Github: https://github.com/weigangq/cov-browser&lt;br /&gt;
File:Rec-fig2.png | Akther, Li, Martin, Di, Sulkow, Pante, Bezrucenlovas, Luft, Qiu (May, 2020). &amp;quot;Origin, recombination, and missed opprotunities:  a genomic perspective of the first 100 days of COVID-19 pandemic&amp;quot;. (Unpublished).&lt;br /&gt;
File:Cov-fig3-trace.png | Akther, Bezrucenlovas, Li, Sulkow, Di, Pante, Martin, Luft, Qiu (Sep, 2021). &amp;quot;Following the Trail of One Million Genomes: Footprints of SARS-CoV-2 Adaptation to Humans&amp;quot;. &#039;&#039;&#039;&#039;&#039;[https://www.biorxiv.org/content/biorxiv/early/2021/05/10/2021.05.07.443114.full.pdf BioRxiv link]&#039;&#039;&#039;&#039;&#039;: . Github:  https://github.com/weigangq/cov-db&lt;br /&gt;
&amp;lt;/gallery&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Lab Resources &amp;amp; Protocols==&lt;br /&gt;
*OneRF usage (Unofficial; login needed): https://sites.google.com/view/onerf/home&lt;br /&gt;
*[https://runestone.academy/ns/books/published/thinkcspy/index.html How to think like a computer scientist: An interactive Python programming book]&lt;br /&gt;
* Nanopore sequencing protocols&lt;br /&gt;
**DNA barcoding: https://nanopore4edu.org/latest/annotated_experiments/dna_barcoding/&lt;br /&gt;
**Yeast genomes: https://nanoporetech.com/document/extraction-method/yeast-dna&lt;br /&gt;
*[[Monte Carlo Club]]&lt;br /&gt;
*[[NY-RaMP Mentoring]]&lt;br /&gt;
*Borreliella genome sequencing consortium: Weekly meetings (Tu @11): Since Jan 2023&lt;br /&gt;
*Borreliella diagnostic antigens (Fall 2023-Fall 2027):&lt;br /&gt;
**Zoom call (Jan 23, 2024)&lt;br /&gt;
** Next meeting: March 23, 2024&lt;br /&gt;
*Qiu lab network [[First Time Guide|first-time user guide]]&lt;br /&gt;
*Qiu lab Github repositories: https://github.com/weigangq/?tab=repositories&lt;br /&gt;
*[[Mini-Tutorals|Mini-Protocols]] (frequently used computer codes and pipelines)&lt;br /&gt;
*Python tutorial: https://wiki.genometracker.org/~weigang/Intro_to_Python.html&lt;br /&gt;
* [[Tick protocol|ick handling protocols]]&lt;br /&gt;
*[[A Primer on the Cluster System at Hunter|Hunter HPC Usage]]&lt;br /&gt;
*[https://r4ds.hadley.nz/ R for Data Science (2e)], (2024) by Wickham, Grolemund &amp;amp; Çetinkaya-Rundel ([https://bookdown.org/ Bookdown version])&lt;br /&gt;
*Borrelia Genome Consortium:  https://www.ncbi.nlm.nih.gov/bioproject/PRJNA431102/&lt;br /&gt;
*Canadian Bbsl genome assemblies: https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA1130942&lt;br /&gt;
* Nanopore sequencing resources:&lt;br /&gt;
**eBook: https://store.nanoporetech.com/us/minion.html&lt;br /&gt;
&lt;br /&gt;
==Wiki Help==&lt;br /&gt;
*[https://www.mediawiki.org/wiki/Special:MyLanguage/Manual:Configuration_settings Configuration settings list]&lt;br /&gt;
*[https://www.mediawiki.org/wiki/Special:MyLanguage/Manual:FAQ MediaWiki FAQ]&lt;br /&gt;
*[https://lists.wikimedia.org/postorius/lists/mediawiki-announce.lists.wikimedia.org/ MediaWiki release mailing list]&lt;br /&gt;
*[https://www.mediawiki.org/wiki/Special:MyLanguage/Localisation#Translation_resources Localise MediaWiki for your language]&lt;br /&gt;
*[https://www.mediawiki.org/wiki/Special:MyLanguage/Manual:Combating_spam Learn how to combat spam on your wiki]&lt;br /&gt;
*Consult the [[mediawikiwiki:Special:MyLanguage/Help:Contents|User&#039;s Guide]] for information on using the wiki software.&lt;/div&gt;</summary>
		<author><name>Wikiuser</name></author>
	</entry>
</feed>