Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions docs/hpc/04_datasets/01_intro.md
Original file line number Diff line number Diff line change
Expand Up @@ -92,15 +92,15 @@ Please open the ImageNet site, find the terms of use ([http://image-net.org/down

NYU has a subscription to Twitter Decahose - 10% random sample of the realtime Twitter Firehose through a streaming connection

*Data are stored* in GCP cloud (BigQuery) and on HPC clusters Greene and Peel (Parquet format).
*Datasets are stored* in GCP cloud (BigQuery) and on the HPC cluster Greene.

Please contact Megan Brown at [The Center for Social Media & Politics](https://csmapnyu.org/) to get access to data and learn the tools available to work with it.

*On cluster dataset is available under (given that you have permissions)*
- `/scratch/work/twitter_decahose/`

### ProQuest Congressional Record
About data set: [ProQuest Congressional Record](https://guides.nyu.edu/tdm/proquest-congressional-record-tdm-guide)
About data set: [ProQuest Congressional Record](https://guides.nyu.edu/govdocs/congressional#s-lg-box-14137380)

The ProQuest Congressional Record text-as-data collection consists of machine-readable files capturing the full text and a small number of metadata fields for a full run of the Congressional Record between 1789 and 2005. Metadata fields include the date of publication, subjects (for issues for which such information exists in the ProQuest system), and URLs linking the full text to the canonical online record for that issue on the ProQuest Congressional platform. A total of 31,952 issues are available.

Expand Down