Archive for category Journals Publishing

Life-science data standards

Posted by peanutbutter in bioinformatics, data standards, Journals Publishing, open data on January 12, 2007

The full complement of the Human Proteome Organisation (HUPO) Proteomics Standards Initiative (PSI), Minimum Information About a Proteomics Experiment (MIAPE), recommended reporting guidelines are now available for community review on the Nature Biotechnology website.

The manuscripts range from the definition of the MIAPE concept to the individual guidelines themselves which cover, Mass Spectrometry, Mass Spectrometry Informatics, Gel Electrophoresis and Molecular Interaction experiments. A further paper on the PSI protein modification ontology (PSI-MOD) is also listed.

Several other Minimum reporting requirements are also listed from other domains such as genome sequences (MIGS) and in Situ Hybridization and Immunochemistry (MISFISHIE).

There is also a paper on the Functional Genomics Experiment Model (FUGE) which is an “extensible framework” or data model for standards in functional genomics, although equally extensible to most scientific experiments.

As this process of community review, hosted on the Nature Biotech website, is a relatively new process and open to anyone, both identified or anonymous, I would encourage anybody with the relevant knowledge to comment on the papers. A greater response from the community ultimately means the guidelines are actually representative of the domain and technology they represent.

Open data

Posted by peanutbutter in bioinformatics, Journals Publishing, open data on November 23, 2006

Open data is a concept I came across while attending the 2nd International Digital Curation Conference and felt it deserved a post in its own right, rather than subsumed by the conference report. I am an advocate of Open Access and feel that Open Data must be a part of this process. What is the point of being able to freely read the publication if you cant freely access the data the publication refers to?

The concept of Open Data was presented to me by Peter Murray-Rust, via the DCC conference, who regularly blogs about the subject. There is a Wikipedia entry which defines the concept and a mailing list which promotes discussion of Open Data.

Steps are underway in the bio-science domain to define Minimum reporting requirements of data for repositories and publications. Two prominent examples are the MGED community with MIAME and the Proteomics community with MAIPE. Within these noble efforts there is no mention of Open Data although it seems the next logical step in the data curation pipeline;

record the defined minimum information and metadata.
structure and present the data.
allow access of the data

Maybe when the effort is made to properly record, structure and describe the data, as these minimum reporting requirements advocate, the scientist and journals will be only too happy to take the next step and declare it Open Data, for the sake of scientific knowledge and progression.

1 Comment

2nd International Digital Curation Conference

Posted by peanutbutter in bioinformatics, conference report, Journals Publishing, open data on November 23, 2006

The Data Curation Center has just held their 2nd International Digital Curation Conference, in Glasgow. The official DCC blog for the conference tracks the thoughts and discourse over the two days and the full program can be found here. As the conference name suggests, the meeting has the particular focus on different aspects of the digital curation life-cycle including managing repositories, educating data scientists and understanding the role of policy and strategy.

One particular talk I was interested in was “The Roles of Shared Data Collections in Neuroscience”. This was presented by a social scientist, as the results of communications with Neuroscientists. Ironically the shared data collection was called “NeuroAnatomical Cell Repository” a pseudonym to “protect the confidentiality of the participants”, so much for the “shared” component! The general conclusions re-iterated what is already know in the bio-sciences; that more experiments are producing large volumes of heterogeneous data that need to be stored, preserved and presented in a manner that allows the efficient use and re-use of the data. There was particular mention that Neuroscience doesn’t have any data reporting standards, a particular buzz-topic in biological sciences.

As a result of this talk, the issue of how we publish this data was again raised, with the provoking statement from the floor, “we should bypass the traditional journals and publish the data ourselves” (a summation of the statement, not an actual quote). This is an issue I have been hearing more and more at recent conferences and in general discussions, a topic that appears to be gathering momentum. Some discourse has already been presented within this blog on some of these issues.

The open panel session on day two, engaged some interesting discussion and I heard a term which I had never heard before, “Open Data“, put forward by Peter Murray-Rust (University of Cambridge). We have all heard of Open Access publishing, (and should not be publishing any other way), but to date this means open access to the journal publication and not the the data that the publication refers to. In something as simple as a graph in a journal publication, generally the access to the numbers/values, has to be re-calculated via a print-out and a ruler. It would be so much easier (and logical) for re-use, analysis or even review, if the presented image was accompanied by the data (even if this was in an excel spreadsheet).

So in summation, the conference presented numerous issues for consideration by a “data scientist” (this may well be the new name for bioinformaticians). The concept of digital data curation is something that is becoming more prevalent in the life-sciences both at the level of the bench scientist (generating metadata) and the analysis, presentation and preservation of the resulting data. No doubt conferences like the DCC will continue to grow in stature and the issues will be further presented in their newly launched International Journal of Digital Curation.

2 Comments

Pubmed search engine

Posted by peanutbutter in bioinformatics, Journals Publishing on November 15, 2006

I cant lay claim that I discovered this all on my own, but as mentioned on Dan Swan’s blog PubMed can be added as a search engine in Firefox 2.0. I checked this out for Hubmed too, a much nicer interface to PubMed, with an RSS feed for search query’s and it works very well.

Journals but not as we know it

Posted by peanutbutter in bioinformatics, Journals Publishing on June 1, 2006

There is an interview on Wired magazine on Harold Varmus, the Nobel laureate and former director of the National Institute of Health. The premise of the article is not on his prize winning exploits but rather his involvement and the setting up of PLoS (The Public Library of Science), which is an open access scientific publishing platform. Many comments and opinions exist on PLoS and its impact but I think the most interesting section of the article is on the new PLoS journal which is due to be launched in the summer called PLoS ONE.

"And this summer, Varmus and his colleagues will launch PLoS One, a paperless journal that will publish online any paper that evaluators deem “scientifically legitimate.” Each article will generate a thread for comment and review. Great papers will be recognized by the discussion they generate, and bad ones will fade away. "Our mission is to transform how science publishing is done,” Varmus says. “We aren’t trying to torpedo the industry. But we are definitely going to change it.” "

This effectively destroys the current peer review system and allows for continuous peer review or an article as more knowledge and opinions develop over the years. It will also permit comment by "experts" that would not have been involved in the traditional peer-review process.

Although there also could be a dark side to the comments where "my paper is better than your paper", or biased comments where a competing interest is evident.

All in all however, I think this is a massive step forward and a massive shake up of the scientific publication process. I hope there is a strong emphasis on providing all the available data in the appropriate standard formats for the domain and if this is the case then I am all for it.

fgibson.com