[IGBF-3047] Investigate: New splice variant annotations for tomato - JIRA UNCC

Details

Type: Task
Status: Closed (View Workflow)
Priority: Major
Resolution: Done
Affects Version/s: None
Fix Version/s: None
Labels:
None

Story Points:
2
Epic Link:
Support NSF pollen grant
Sprint:
Spring 9 2022 May 9, Summer 1 2022 May 23, Summer 2 2022 June 6, Summer 3 2022 June 21, Summer 4 2022 July 4

Description

The S. lycopersicon (cultivated tomato) gene annotations include only one gene model per gene. However, visualizing RNA-Seq data in IGB shows that a large number of genes produce multiple splice forms. At least one other group has noticed this, as well. In their article "Expanding Alternative Splicing Identification by Integrating Multiple Sources of Transcription Data in Tomato", a group at Ohio State University led by Prof. Xiangjia (Jack) Min reported using transcriptome data, including ESTs and RNA-Seq data, to assemble new gene models. I downloaded these and deployed them to IGB Quickload; they are one of the available data sets for the next to last genome release.

There may be other groups developing similar datasets for the most recent genome release for tomato. And in order to quantify splice variant expression using current methods, it would be extremely helpful to have an up-to-date, accurate-as-possible collection of gene models annotated with functional information. Who else is interested in this and would be interested in contributing? Or, is this something only our group might care about?

As part of the pollen NSF project, we are trying to understand and discover how heat stress triggers changes in RNA synthesis in pollen, in pollen tubes, and in other sample types related to reproduction in plants, especially tomato?

How homogenous are the RNA-Seq data sets coming from the pollen project? So far, all the data have been from a single cell type: germinating pollen tubes. I do not recall seeing much evidence for alternative splicing in these datasets, at least not as compared with other samples that included many cell types, e.g., root or shoot. Also, are there splice forms that exist mainly in pollen but not other tissue types? We found some examples of this in the Arabidospis pollen RNA-Seq data described in our paper "RNA-seq of Arabidopsis pollen uncovers novel transcription and alternative splicing".

How many tomato RNA-Seq data sets are there, and how good are they? For the purpose of producing new gene models, the best bulk RNA-Seq data would be paired end, very long read lengths, and strand-specific. Are such data available currently, or would we need to create new data to cover the entirety of transcription?

Attachments

Options
- Sort By Name
- Sort By Date
- Ascending
- Descending
- Thumbnails
- List
- Download All

Attachments

SRP328042 _tomato_Drought_stress.txt
0.2 kB
22/Jun/22 12:33 PM

Issue Links

relates to

IGBF-3135 Add new tomato genome and annotations to IGB Quickload repository

Closed

Activity

No work has yet been logged on this issue.

People

Assignee:

Ann Loraine

Reporter:

Ann Loraine

Votes:

0 Vote for this issue

Watchers:

2 Start watching this issue

Dates

Created:

04/Jan/22 11:49 PM

Updated:

07/Jul/22 5:48 PM

Resolved:

07/Jul/22 5:48 PM