Getting Unremediated PDB Files: Difference between revisions

From Proteopedia
Jump to navigationJump to search
Angel Herraez (talk | contribs)
7-zip as a possible uncompressor
Eric Martz (talk | contribs)
 
(8 intermediate revisions by the same user not shown)
Line 1: Line 1:
Two rounds of remediation have now taken place between 2007 and 2009 to better standardize and enhance the PDB files deposited before December 2, 2008. The details of the two rounds can be found at [http://www.wwpdb.org/docs.html the Worldwide Protein Data Bank's Documentation page]. The most recent version was released March 17th, 2009 as described on the [http://www.pdb.org/pdb/static.do?p=general_information/news_publications/news/news_2009.html#20090317a news page at the Protein Data Bank].
Periodically, the [[Protein Data Bank]] remediates [[PDB files]] in the worldwide archive that it maintains. Two rounds of [[PDB#Remediation|remediation]] took place between 2007 and 2009, in order to better standardize and enhance the PDB files deposited before December 2, 2008. The details of the two rounds can be found at [http://www.wwpdb.org/docs.html the Worldwide Protein Data Bank's Documentation page]. The most recent version was released March 17th, 2009 as described on the [http://www.pdb.org/pdb/static.do?p=general_information/news_publications/news/news_2009.html#20090317a news page at the Protein Data Bank].


 
The unremediated PDB archive from before March 17, 2009 is available, as detailed [http://www.pdb.org/pdb/static.do?p=general_information/news_publications/news/news_2009.html#20090317a here] because a time-stamped snapshot of the PDB archive before the March 17th release exists [ftp://snapshots.wwpdb.org/ here] in the directory 20090316.  
The unremediated PDB archive from before March 17, 2008 is available, as detailed [http://www.pdb.org/pdb/static.do?p=general_information/news_publications/news/news_2009.html#20090317a here] because a time-stamped snapshot of the PDB archive before the March 17th release exists [ftp://snapshots.wwpdb.org/ here] in the directory 20090316.  


The unremediated PDB archive from before August 1, 2007 is available, as detailed [http://www.rcsb.org/pdb/static.do?p=general_information/news_publications/news/news_2007.html#20070904 here].
The unremediated PDB archive from before August 1, 2007 is available, as detailed [http://www.rcsb.org/pdb/static.do?p=general_information/news_publications/news/news_2007.html#20070904 here].
Line 8: Line 7:
If a PDB file was released after December 2, 2008, it is not available in unremediated form.
If a PDB file was released after December 2, 2008, it is not available in unremediated form.


==Proteopedia avoids remediation-related problems==


===Why would you need an unremediated version of a pdb file?===
Following the experience of the 2009 remediation, Proteopedia automatically saves the version of the PDB file for which each molecular scene is developed, along with the Jmol script for the scene. When a new scene is developed with the [[SAT|Scene Authoring Tools]], the current version of the PDB file is used and saved. Thus, subsequent remediations cannot inadvertantly corrupt scenes developed on earlier versions of the PDB file. (The 2009 remediation changed the order of atoms in some PDB files, which broke a few scenes until Proteopedia was modified to save the PDB file along with the scene script. These were repaired by obtaining the unremediated PDB files and using them for these scenes.)
A significant change made in the course of the first round (2007) of remediation was the distinction between ribonucleotides (A, C, G, I, T, U) and deoxyribonucleotides (DA, DC, DG, DI, DT, DU). The main reason for getting unremediated PDB files from before the 2007 remediation is that when the remediated PDB files contain DNA, CHIME-based Protein Explorer (and perhaps some other software) does not display the DNA properly. If the PDB file does not contain DNA (protein, RNA, solvent and ligands are OK), you probably don't need the unremediated file. If a PDB file was released after August 1, 2007, it will not be available in unremediated form that suits CHIME-based (and perhaps other) software. The second round of remediation (2008 round; released March 17th 2000) also mainly affected nucleic acid residues and atoms.<br>
Proteopedia also needs the unremediated files (pre-March 17th 2009). Proteopedia actually premiered between the two rounds of remediation and relied on the atom serial numbers for saving the scenes, yet contacts the PDB to get the current PDB file. Thus when the March 17th remediated version of the database was released with differing atom serial numbers, scenes involving nucleic acid using the newer files often no longer looked correct because some atom serial numbers now do not match. A global fix for this is currently being worked on according to Eran Hodis and Jaime Prilusky (see the gray banner at the top of the page for updates).


===How to get the unremediated version?===
==Why would you need an unremediated version of a pdb file?==
Use the simple interface [http://www.umass.edu/microbio/chime/pe_beta/pe/protexpl/unremed.htm here at Eric Martz's UMASS site] to easily get July 31, 2007 unremediated pdb files [ftp://snapshots.wwpdb.org/ via ftp at the RCSB Protein Data Bank] in the directory 20070731. This is primarily for obtaining DNA before the residues were re-named DC, DG, DT, DA, for Protein Explorer/Chime.
A significant change made in the course of the first round (2007) of remediation was the distinction between ribonucleotides (A, C, G, I, T, U) and deoxyribonucleotides (DA, DC, DG, DI, DT, DU). The main reason for getting unremediated PDB files from before the 2007 remediation is that when the remediated PDB files contain DNA, [[Chime]]-based [[Protein Explorer]] (and perhaps some other software) does not display the DNA properly. If the PDB file does not contain DNA (protein, RNA, solvent and ligands are OK), you probably don't need the unremediated file. If a PDB file was released after August 1, 2007, it will not be available in unremediated form that suits CHIME-based (and perhaps other) software. The second round of remediation (2009 round; released March 17th 2009) also mainly affected nucleic acid residues and atoms.


The March 16, 2009 unremediated versions are available [ftp://snapshots.wwpdb.org/ via ftp at the RCSB Protein Data Bank] in the  
==How to get the unremediated version?==
<!--Use the simple interface [http://www.umass.edu/microbio/chime/pe_beta/pe/protexpl/unremed.htm here at Eric Martz's UMASS site] to easily get July 31, 2007 unremediated pdb files [ftp://snapshots.wwpdb.org/ via ftp at the RCSB Protein Data Bank] in the directory 20070731. This is primarily for obtaining DNA before the residues were re-named DC, DG, DT, DA, for Protein Explorer/Chime.-->
 
The March 16, 2009 unremediated versions are available [ftp://snapshots.wwpdb.org/ via ftp at the World Wide Protein Data Bank] in the  
directory 20090316.
directory 20090316.


Line 26: Line 27:


To uncompress the downloaded files:
To uncompress the downloaded files:
* First try simply double-clicking the compressed file. On Macs (in 2019) this decompresses them.
* For uncompressing the 2007 files that are in .Z format:
* For uncompressing the 2007 files that are in .Z format:
** Both .Z and .gz files can be uncompressed in Windows using [http://www.7-zip.org/ 7-Zip], which is freeware and open source.
** Both .Z and .gz files can be uncompressed in Windows using [http://www.7-zip.org/ 7-Zip], which is freeware and open source.
Line 38: Line 40:
''Please, note that .gz files can be displayed in Proteopedia without being uncompressed, since Jmol can read gzipped files directly.''
''Please, note that .gz files can be displayed in Proteopedia without being uncompressed, since Jmol can read gzipped files directly.''


 
==See Also==
See also [[Standard Residues]] and [[Non-Standard Residues]]
*[[Atomic coordinate files]]
*[[PDB file format]]
*[[PDB#Remediation|Remediation]]
*[[Standard Residues]]
*[[Non-Standard Residues]]