Uploaded image for project: 'IGB'
  1. IGB
  2. IGBF-3244

Run rnaseq pipeline on mark-2022-timeseries

    Details

    • Type: Task
    • Status: Closed (View Workflow)
    • Priority: Major
    • Resolution: Done
    • Affects Version/s: None
    • Fix Version/s: None
    • Labels:
      None

      Description

      Run nextflow for the dataset in:

      /projects/tomato_genome/rnaseq/mark-2022-timeseries/30-771363348/00_fastq

      This is the "time course" dataset discussed by Rasha at the 2023-01-17 group meeting. Note that she has already run nextflow for this dataset but using "unstranded" for the "strandedness" parameter in the "samples.csv" file. It turns out this dataset comes from libraries that were created using a strand-specific RNA-Seq library. To be on the safe side, we should re-run the pipeline using parameter "reverse", as indicated in the multiQC report included with Rasha's initial run of the nextflow nf-core rnaseq pipeline.

      Kindly run the nf-core pipeline in this location:

      • /nobackup/tomato_genome/mark-2022-timeseries

      Note on attached files:

      • multiqc report on the entire run done by Rasha is attached, copied from google drive location GTTR-NSF PGRP - 2020-24 IOS-1939255 > Experiments > Rasha_RNA-seq_Time_Course > Results > multiqc > star_salmon > multiqc_report.html
      • Link: https://drive.google.com/drive/u/1/folders/1GJnZefP-7TE-ch-c0lblGZMOqpwSgZRK
      • 2023-01-18_timeseries_multiqc_report.html - MultiQC report from re-running nextflow (Molly's new work)\
      • sample.csv - new samples file used to re-run nextflow (Molly's new work)

        Attachments

          Issue Links

            Activity

            aloraine Ann Loraine created issue -
            aloraine Ann Loraine made changes -
            Field Original Value New Value
            Epic Link IGBF-2993 [ 21429 ]
            aloraine Ann Loraine made changes -
            Assignee Ann Loraine [ aloraine ]
            aloraine Ann Loraine made changes -
            Summary Re-run nfcore rnaseq pipeline for seedlingPollen data Run nfcore rnaseq pipeline for time course data
            aloraine Ann Loraine made changes -
            Description We need to re-run nf-core pipeline with a new strandedness parameter as noted in attached multiqc report.
            aloraine Ann Loraine made changes -
            Description Run nextflow for the dataset in:

            /projects/tomato_genome/rnaseq/mark-2022-timeseries/30-771363348/00_fastq

            This is the "time course" dataset discussed by Rasha at the 2023-01-17 group meeting. Note that she has already run nextflow for this dataset but using "unstranded" for the "strandedness" parameter in the "samples.csv" file. It turns out this dataset comes from libraries that were created using a strand-specific RNA-Seq library. To be on the safe side, we should re-run the pipeline using parameter "reverse", as indicated in the multiQC report included with Rasha's initial run of the nextflow nf-core rnaseq pipeline.

            Run the nf-core pipeline in this location: /

            * nobackup/tomato_genome/mark-2022-timeseries
            aloraine Ann Loraine made changes -
            Summary Run nfcore rnaseq pipeline for time course data Run rnaseq pipeline on
            aloraine Ann Loraine made changes -
            Summary Run rnaseq pipeline on Run rnaseq pipeline on mark-2022-timeseries
            aloraine Ann Loraine made changes -
            Attachment 2022-11-24_multiqc_report.html [ 17651 ]
            aloraine Ann Loraine made changes -
            Description Run nextflow for the dataset in:

            /projects/tomato_genome/rnaseq/mark-2022-timeseries/30-771363348/00_fastq

            This is the "time course" dataset discussed by Rasha at the 2023-01-17 group meeting. Note that she has already run nextflow for this dataset but using "unstranded" for the "strandedness" parameter in the "samples.csv" file. It turns out this dataset comes from libraries that were created using a strand-specific RNA-Seq library. To be on the safe side, we should re-run the pipeline using parameter "reverse", as indicated in the multiQC report included with Rasha's initial run of the nextflow nf-core rnaseq pipeline.

            Run the nf-core pipeline in this location: /

            * nobackup/tomato_genome/mark-2022-timeseries
            Run nextflow for the dataset in:

            /projects/tomato_genome/rnaseq/mark-2022-timeseries/30-771363348/00_fastq

            This is the "time course" dataset discussed by Rasha at the 2023-01-17 group meeting. Note that she has already run nextflow for this dataset but using "unstranded" for the "strandedness" parameter in the "samples.csv" file. It turns out this dataset comes from libraries that were created using a strand-specific RNA-Seq library. To be on the safe side, we should re-run the pipeline using parameter "reverse", as indicated in the multiQC report included with Rasha's initial run of the nextflow nf-core rnaseq pipeline.

            Run the nf-core pipeline in this location: /

            * nobackup/tomato_genome/mark-2022-timeseries

            Note on attachments:

            * multiqc report on the entire run done by Rasha is attached, copied from google drive location GTTR-NSF PGRP - 2020-24 IOS-1939255 > Experiments > Rasha_RNA-seq_Time_Course > Results > multiqc > star_salmon > multiqc_report.html
            aloraine Ann Loraine made changes -
            Description Run nextflow for the dataset in:

            /projects/tomato_genome/rnaseq/mark-2022-timeseries/30-771363348/00_fastq

            This is the "time course" dataset discussed by Rasha at the 2023-01-17 group meeting. Note that she has already run nextflow for this dataset but using "unstranded" for the "strandedness" parameter in the "samples.csv" file. It turns out this dataset comes from libraries that were created using a strand-specific RNA-Seq library. To be on the safe side, we should re-run the pipeline using parameter "reverse", as indicated in the multiQC report included with Rasha's initial run of the nextflow nf-core rnaseq pipeline.

            Run the nf-core pipeline in this location: /

            * nobackup/tomato_genome/mark-2022-timeseries

            Note on attachments:

            * multiqc report on the entire run done by Rasha is attached, copied from google drive location GTTR-NSF PGRP - 2020-24 IOS-1939255 > Experiments > Rasha_RNA-seq_Time_Course > Results > multiqc > star_salmon > multiqc_report.html
            Run nextflow for the dataset in:

            /projects/tomato_genome/rnaseq/mark-2022-timeseries/30-771363348/00_fastq

            This is the "time course" dataset discussed by Rasha at the 2023-01-17 group meeting. Note that she has already run nextflow for this dataset but using "unstranded" for the "strandedness" parameter in the "samples.csv" file. It turns out this dataset comes from libraries that were created using a strand-specific RNA-Seq library. To be on the safe side, we should re-run the pipeline using parameter "reverse", as indicated in the multiQC report included with Rasha's initial run of the nextflow nf-core rnaseq pipeline.

            Kindly run the nf-core pipeline in this location: /

            * nobackup/tomato_genome/mark-2022-timeseries

            Note on attached files:

            * multiqc report on the entire run done by Rasha is attached, copied from google drive location GTTR-NSF PGRP - 2020-24 IOS-1939255 > Experiments > Rasha_RNA-seq_Time_Course > Results > multiqc > star_salmon > multiqc_report.html
            aloraine Ann Loraine made changes -
            Description Run nextflow for the dataset in:

            /projects/tomato_genome/rnaseq/mark-2022-timeseries/30-771363348/00_fastq

            This is the "time course" dataset discussed by Rasha at the 2023-01-17 group meeting. Note that she has already run nextflow for this dataset but using "unstranded" for the "strandedness" parameter in the "samples.csv" file. It turns out this dataset comes from libraries that were created using a strand-specific RNA-Seq library. To be on the safe side, we should re-run the pipeline using parameter "reverse", as indicated in the multiQC report included with Rasha's initial run of the nextflow nf-core rnaseq pipeline.

            Kindly run the nf-core pipeline in this location: /

            * nobackup/tomato_genome/mark-2022-timeseries

            Note on attached files:

            * multiqc report on the entire run done by Rasha is attached, copied from google drive location GTTR-NSF PGRP - 2020-24 IOS-1939255 > Experiments > Rasha_RNA-seq_Time_Course > Results > multiqc > star_salmon > multiqc_report.html
            Run nextflow for the dataset in:

            /projects/tomato_genome/rnaseq/mark-2022-timeseries/30-771363348/00_fastq

            This is the "time course" dataset discussed by Rasha at the 2023-01-17 group meeting. Note that she has already run nextflow for this dataset but using "unstranded" for the "strandedness" parameter in the "samples.csv" file. It turns out this dataset comes from libraries that were created using a strand-specific RNA-Seq library. To be on the safe side, we should re-run the pipeline using parameter "reverse", as indicated in the multiQC report included with Rasha's initial run of the nextflow nf-core rnaseq pipeline.

            Kindly run the nf-core pipeline in this location:

            * /nobackup/tomato_genome/mark-2022-timeseries

            Note on attached files:

            * multiqc report on the entire run done by Rasha is attached, copied from google drive location GTTR-NSF PGRP - 2020-24 IOS-1939255 > Experiments > Rasha_RNA-seq_Time_Course > Results > multiqc > star_salmon > multiqc_report.html
            Mdavis4290 Molly Davis (Inactive) made changes -
            Assignee Molly Davis [ molly ]
            Mdavis4290 Molly Davis (Inactive) made changes -
            Description Run nextflow for the dataset in:

            /projects/tomato_genome/rnaseq/mark-2022-timeseries/30-771363348/00_fastq

            This is the "time course" dataset discussed by Rasha at the 2023-01-17 group meeting. Note that she has already run nextflow for this dataset but using "unstranded" for the "strandedness" parameter in the "samples.csv" file. It turns out this dataset comes from libraries that were created using a strand-specific RNA-Seq library. To be on the safe side, we should re-run the pipeline using parameter "reverse", as indicated in the multiQC report included with Rasha's initial run of the nextflow nf-core rnaseq pipeline.

            Kindly run the nf-core pipeline in this location:

            * /nobackup/tomato_genome/mark-2022-timeseries

            Note on attached files:

            * multiqc report on the entire run done by Rasha is attached, copied from google drive location GTTR-NSF PGRP - 2020-24 IOS-1939255 > Experiments > Rasha_RNA-seq_Time_Course > Results > multiqc > star_salmon > multiqc_report.html
            Run nextflow for the dataset in:

            /projects/tomato_genome/rnaseq/mark-2022-timeseries/30-771363348/00_fastq

            This is the "time course" dataset discussed by Rasha at the 2023-01-17 group meeting. Note that she has already run nextflow for this dataset but using "unstranded" for the "strandedness" parameter in the "samples.csv" file. It turns out this dataset comes from libraries that were created using a strand-specific RNA-Seq library. To be on the safe side, we should re-run the pipeline using parameter "reverse", as indicated in the multiQC report included with Rasha's initial run of the nextflow nf-core rnaseq pipeline.

            Kindly run the nf-core pipeline in this location:

            * /nobackup/tomato_genome/mark-2022-timeseries

            Note on attached files:

            * multiqc report on the entire run done by Rasha is attached, copied from google drive location GTTR-NSF PGRP - 2020-24 IOS-1939255 > Experiments > Rasha_RNA-seq_Time_Course > Results > multiqc > star_salmon > multiqc_report.html
            * Link: https://drive.google.com/drive/u/1/folders/1GJnZefP-7TE-ch-c0lblGZMOqpwSgZRK
            Hide
            Mdavis4290 Molly Davis (Inactive) added a comment - - edited

            Next steps: Find or make a csv sample sheet and change strandedness to 'reverse' to run nextflow.
            sample.csv

            Show
            Mdavis4290 Molly Davis (Inactive) added a comment - - edited Next steps: Find or make a csv sample sheet and change strandedness to 'reverse' to run nextflow. sample.csv
            Mdavis4290 Molly Davis (Inactive) made changes -
            Attachment sample.csv [ 17652 ]
            Mdavis4290 Molly Davis (Inactive) made changes -
            Status To-Do [ 10305 ] In Progress [ 3 ]
            Mdavis4290 Molly Davis (Inactive) made changes -
            Hide
            Mdavis4290 Molly Davis (Inactive) added a comment - - edited


            Nextflow Pipeline Ran Successfully!
            Directory: /nobackup/tomato_genome/mark-2022-timeseries
            Next steps: Rename sorted bam files and make scaled coverage graphs.

            Show
            Mdavis4290 Molly Davis (Inactive) added a comment - - edited Nextflow Pipeline Ran Successfully! Directory: /nobackup/tomato_genome/mark-2022-timeseries Next steps: Rename sorted bam files and make scaled coverage graphs.
            Hide
            Mdavis4290 Molly Davis (Inactive) added a comment - - edited

            Scaled coverage graphs have been made and are located:

            /nobackup/tomato_genome/mark-2022-timeseries/results/star_salmon
            

            Notes: I can move the coverage graphs to their own directory if you would like. Let me know!

            Multiqc report:

            scp mdavi258@hpc.uncc.edu:/nobackup/tomato_genome/mark-2022-timeseries/results/multiqc/star_salmon/multiqc_report.html timeseries_multiqc_report.html
            

            [^timeseries_multiqc_report.html]

            Notes: Multiqc report seems to show better mapping and correct strandedness now compared to the previous report and nextflow run.

            Next step: Pipeline, coverage graphs, and Multiqc report need to be reviewed.
            Ann Loraine

            Show
            Mdavis4290 Molly Davis (Inactive) added a comment - - edited Scaled coverage graphs have been made and are located: /nobackup/tomato_genome/mark-2022-timeseries/results/star_salmon Notes: I can move the coverage graphs to their own directory if you would like. Let me know! Multiqc report: scp mdavi258@hpc.uncc.edu:/nobackup/tomato_genome/mark-2022-timeseries/results/multiqc/star_salmon/multiqc_report.html timeseries_multiqc_report.html [^timeseries_multiqc_report.html] Notes: Multiqc report seems to show better mapping and correct strandedness now compared to the previous report and nextflow run. Next step : Pipeline, coverage graphs, and Multiqc report need to be reviewed. Ann Loraine
            Mdavis4290 Molly Davis (Inactive) made changes -
            Assignee Molly Davis [ molly ]
            Mdavis4290 Molly Davis (Inactive) made changes -
            Status In Progress [ 3 ] Needs 1st Level Review [ 10005 ]
            Mdavis4290 Molly Davis (Inactive) made changes -
            Attachment timeseries_multiqc_report.html [ 17658 ]
            aloraine Ann Loraine made changes -
            Assignee Ann Loraine [ aloraine ]
            aloraine Ann Loraine made changes -
            Status Needs 1st Level Review [ 10005 ] First Level Review in Progress [ 10301 ]
            Hide
            aloraine Ann Loraine added a comment -

            I reviewed multiqc report and noticed no problems.
            I migrated coverage graphs and bam files to igb quickload host and updated makeAnnotsXml.py in https://bitbucket.org/hotpollen/splicing-analysis/src/main/ to use the new files.
            See: ManageQuickload/makeAnnotsXml.py and ManageQuickload/quickload/S_lycopersicum_Jun_2022/annots.xml.

            Moving to DONE.

            Show
            aloraine Ann Loraine added a comment - I reviewed multiqc report and noticed no problems. I migrated coverage graphs and bam files to igb quickload host and updated makeAnnotsXml.py in https://bitbucket.org/hotpollen/splicing-analysis/src/main/ to use the new files. See: ManageQuickload/makeAnnotsXml.py and ManageQuickload/quickload/S_lycopersicum_Jun_2022/annots.xml. Moving to DONE.
            aloraine Ann Loraine made changes -
            Status First Level Review in Progress [ 10301 ] Needs 1st Level Review [ 10005 ]
            aloraine Ann Loraine made changes -
            Status Needs 1st Level Review [ 10005 ] First Level Review in Progress [ 10301 ]
            aloraine Ann Loraine made changes -
            Status First Level Review in Progress [ 10301 ] Ready for Pull Request [ 10304 ]
            aloraine Ann Loraine made changes -
            Status Ready for Pull Request [ 10304 ] Pull Request Submitted [ 10101 ]
            aloraine Ann Loraine made changes -
            Status Pull Request Submitted [ 10101 ] Reviewing Pull Request [ 10303 ]
            aloraine Ann Loraine made changes -
            Status Reviewing Pull Request [ 10303 ] Merged Needs Testing [ 10002 ]
            aloraine Ann Loraine made changes -
            Status Merged Needs Testing [ 10002 ] Post-merge Testing In Progress [ 10003 ]
            aloraine Ann Loraine made changes -
            Resolution Done [ 10000 ]
            Status Post-merge Testing In Progress [ 10003 ] Closed [ 6 ]
            aloraine Ann Loraine made changes -
            Assignee Ann Loraine [ aloraine ] Molly Davis [ molly ]
            Hide
            aloraine Ann Loraine added a comment - - edited

            I noticed that coverage graphs for this new dataset, which is strand-specific and paired-end, look a bit different, with different patterns of peaks and valleys, compared to earlier data from Genewiz where the data were paired-end and NOT strand-specific. Weird. I don't know why this occurred.

            For example, see:

            GenomeBrowserImages/TimeCourseVsOlderData-CoverageGraphProfileDifference.png

            Creating new ticket to investigate.

            Show
            aloraine Ann Loraine added a comment - - edited I noticed that coverage graphs for this new dataset, which is strand-specific and paired-end, look a bit different, with different patterns of peaks and valleys, compared to earlier data from Genewiz where the data were paired-end and NOT strand-specific. Weird. I don't know why this occurred. For example, see: GenomeBrowserImages/TimeCourseVsOlderData-CoverageGraphProfileDifference.png Creating new ticket to investigate.
            aloraine Ann Loraine made changes -
            Link This issue relates to IGBF-3245 [ IGBF-3245 ]
            aloraine Ann Loraine made changes -
            Attachment timeseries_multiqc_report.html [ 17658 ]
            aloraine Ann Loraine made changes -
            aloraine Ann Loraine made changes -
            Description Run nextflow for the dataset in:

            /projects/tomato_genome/rnaseq/mark-2022-timeseries/30-771363348/00_fastq

            This is the "time course" dataset discussed by Rasha at the 2023-01-17 group meeting. Note that she has already run nextflow for this dataset but using "unstranded" for the "strandedness" parameter in the "samples.csv" file. It turns out this dataset comes from libraries that were created using a strand-specific RNA-Seq library. To be on the safe side, we should re-run the pipeline using parameter "reverse", as indicated in the multiQC report included with Rasha's initial run of the nextflow nf-core rnaseq pipeline.

            Kindly run the nf-core pipeline in this location:

            * /nobackup/tomato_genome/mark-2022-timeseries

            Note on attached files:

            * multiqc report on the entire run done by Rasha is attached, copied from google drive location GTTR-NSF PGRP - 2020-24 IOS-1939255 > Experiments > Rasha_RNA-seq_Time_Course > Results > multiqc > star_salmon > multiqc_report.html
            * Link: https://drive.google.com/drive/u/1/folders/1GJnZefP-7TE-ch-c0lblGZMOqpwSgZRK
            Run nextflow for the dataset in:

            /projects/tomato_genome/rnaseq/mark-2022-timeseries/30-771363348/00_fastq

            This is the "time course" dataset discussed by Rasha at the 2023-01-17 group meeting. Note that she has already run nextflow for this dataset but using "unstranded" for the "strandedness" parameter in the "samples.csv" file. It turns out this dataset comes from libraries that were created using a strand-specific RNA-Seq library. To be on the safe side, we should re-run the pipeline using parameter "reverse", as indicated in the multiQC report included with Rasha's initial run of the nextflow nf-core rnaseq pipeline.

            Kindly run the nf-core pipeline in this location:

            * /nobackup/tomato_genome/mark-2022-timeseries

            Note on attached files:

            * multiqc report on the entire run done by Rasha is attached, copied from google drive location GTTR-NSF PGRP - 2020-24 IOS-1939255 > Experiments > Rasha_RNA-seq_Time_Course > Results > multiqc > star_salmon > multiqc_report.html
            * Link: https://drive.google.com/drive/u/1/folders/1GJnZefP-7TE-ch-c0lblGZMOqpwSgZRK
            * 2023-01-18_timeseries_multiqc_report.html - MultiQC report from re-running nextflow (Molly's new work)
            aloraine Ann Loraine made changes -
            Description Run nextflow for the dataset in:

            /projects/tomato_genome/rnaseq/mark-2022-timeseries/30-771363348/00_fastq

            This is the "time course" dataset discussed by Rasha at the 2023-01-17 group meeting. Note that she has already run nextflow for this dataset but using "unstranded" for the "strandedness" parameter in the "samples.csv" file. It turns out this dataset comes from libraries that were created using a strand-specific RNA-Seq library. To be on the safe side, we should re-run the pipeline using parameter "reverse", as indicated in the multiQC report included with Rasha's initial run of the nextflow nf-core rnaseq pipeline.

            Kindly run the nf-core pipeline in this location:

            * /nobackup/tomato_genome/mark-2022-timeseries

            Note on attached files:

            * multiqc report on the entire run done by Rasha is attached, copied from google drive location GTTR-NSF PGRP - 2020-24 IOS-1939255 > Experiments > Rasha_RNA-seq_Time_Course > Results > multiqc > star_salmon > multiqc_report.html
            * Link: https://drive.google.com/drive/u/1/folders/1GJnZefP-7TE-ch-c0lblGZMOqpwSgZRK
            * 2023-01-18_timeseries_multiqc_report.html - MultiQC report from re-running nextflow (Molly's new work)
            Run nextflow for the dataset in:

            /projects/tomato_genome/rnaseq/mark-2022-timeseries/30-771363348/00_fastq

            This is the "time course" dataset discussed by Rasha at the 2023-01-17 group meeting. Note that she has already run nextflow for this dataset but using "unstranded" for the "strandedness" parameter in the "samples.csv" file. It turns out this dataset comes from libraries that were created using a strand-specific RNA-Seq library. To be on the safe side, we should re-run the pipeline using parameter "reverse", as indicated in the multiQC report included with Rasha's initial run of the nextflow nf-core rnaseq pipeline.

            Kindly run the nf-core pipeline in this location:

            * /nobackup/tomato_genome/mark-2022-timeseries

            Note on attached files:

            * multiqc report on the entire run done by Rasha is attached, copied from google drive location GTTR-NSF PGRP - 2020-24 IOS-1939255 > Experiments > Rasha_RNA-seq_Time_Course > Results > multiqc > star_salmon > multiqc_report.html
            * Link: https://drive.google.com/drive/u/1/folders/1GJnZefP-7TE-ch-c0lblGZMOqpwSgZRK
            * 2023-01-18_timeseries_multiqc_report.html - MultiQC report from re-running nextflow (Molly's new work)\
            * sample.csv - new samples file used to re-run nextflow (Molly's new work)
            aloraine Ann Loraine made changes -
            Link This issue relates to IGBF-3246 [ IGBF-3246 ]

              People

              • Assignee:
                Mdavis4290 Molly Davis (Inactive)
                Reporter:
                aloraine Ann Loraine
              • Votes:
                0 Vote for this issue
                Watchers:
                2 Start watching this issue

                Dates

                • Created:
                  Updated:
                  Resolved: