WARC file for University at Albany - SUNY - Office of Facilities Management, 2016 June 17

Acquisition information:

crawl: 220038

Crawl Rules

Ignore Robots.txt for www.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for www.alumni.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for www.ualbanysports.com (last updated 2016-02-11)

Ignore Robots.txt for library.albany.edu (last updated 2017-05-19)

Ignore Robots.txt for alumni.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for asrc.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for atmos.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for bioinformatics.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for cela.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for choose.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for cs.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for csda.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for imls.ctg.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for www.ctg.albany.edu (last updated 2017-05-19)

Ignore Robots.txt for cwig.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for events.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for hr.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for ibl.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for illiad.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for liblogs.albany.edu (last updated 2017-05-19)

Ignore Robots.txt for libguides.library.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for scholarsarchive.library.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for listserv.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for m.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for math.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for omega.math.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for mumford.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for nyjm.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for pdp.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for resnet.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for rit.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for cyberphysics.rit.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for rna.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for slsc.albany.edu (last updated 2017-05-19)

Ignore Robots.txt for uaems.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for uapps.albany.edu (last updated 2016-02-11)

Ignore Robots.txt for wiki.albany.edu (last updated 2016-02-11)

Crawl Times

start_date: 2016-06-17T16:55:03Z

original_start_date: 2016-06-17T16:55:03Z

last_resumption: None

processing_end_date: 2016-06-23T01:46:32Z

end_date: 2016-06-22T19:28:00Z

elapsed_ms: 441165235

Crawl Types

type: MONTHLY

recurrence_type: MONTHLY

pdfs_only: False

test: False

Crawl Limits

time_limit: 432000

document_limit: None

byte_limit: None

crawl_stop_requested: None

Crawl Results

status: FINISHED_TIME_LIMIT

discovered_count: 8887536

novel_count: 557229

duplicate_count: 1067492

resumption_count: 0

queued_count: 7262815

downloaded_count: 1624721

download_failures: 247

warc_revisit_count: 1067450

warc_url_count: 1624592

total_data_in_kbs: 251943396

duplicate_bytes: 239433481575

warc_compressed_bytes: 1787914405

Crawl Technical Details

doc_rate: 3.68

kb_rate: 571.0

Physical / technical requirements:
Researchers interested in data analysis with web archives may request a WARC file. WARC files are very large and difficult to work with. Your request may take time to process, and we may be unable to deliver your request remotely. Please consult an archivist if you are interested in advanced research with web archives.

Using these materials

Access:
The archives are open to the public and anyone is welcome to visit and view the collections.
Collection restrictions:
Access to this collection is unrestricted.
Collection terms of access:
The researcher assumes full responsibility for conforming to the laws of copyright. Some materials in these collections may be protected by the U.S. Copyright Law (Title 17, U.S.C.) and/or by the copyright or neighboring-rights laws of other nations. More information about U.S. Copyright is provided by the Copyright Office. Additionally, re-use may be restricted by terms of University Libraries gift or purchase agreements, donor restrictions, privacy and publicity rights, licensing and trademarks.

Access options

Ask an Archivist

Ask a question or schedule an individualized meeting to discuss archival materials and potential research needs.

Schedule a Visit

Archival materials can be viewed in-person in our reading room. We recommend making an appointment to ensure materials are available when you arrive.

Make a Remote Request

We may also be able to deliver digital scans remotely for a fee.