<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Nextflow on @micheloosterhof</title><link>https://www.micheloosterhof.com/tags/nextflow/</link><description>Recent content in Nextflow on @micheloosterhof</description><generator>Hugo</generator><language>en</language><lastBuildDate>Mon, 27 Jul 2026 14:44:17 +0800</lastBuildDate><atom:link href="https://www.micheloosterhof.com/tags/nextflow/index.xml" rel="self" type="application/rss+xml"/><item><title>Analyzing my own genome on Google Cloud Batch</title><link>https://www.micheloosterhof.com/posts/2026-07-27-personal-wgs-gcp-batch/</link><pubDate>Mon, 27 Jul 2026 00:00:00 +0000</pubDate><guid>https://www.micheloosterhof.com/posts/2026-07-27-personal-wgs-gcp-batch/</guid><description>&lt;p&gt;I had my genome sequenced, and got back two files: a pair of gzipped
&lt;a href="https://en.wikipedia.org/wiki/FASTQ_format" class="external-link" target="_blank" rel="noopener"&gt;FASTQs&lt;/a&gt;, about
150 GB, holding a few billion short DNA reads. That is the raw output of a sequencer —
essentially a giant pile of 150-letter fragments with no idea where in the genome each
one belongs. Turning that into something you can actually read — &lt;em&gt;here is where you
differ from the reference human genome, and here is what those differences might mean&lt;/em&gt; —
is a real computational pipeline.&lt;/p&gt;</description></item></channel></rss>