Dataset Search

Displaying 351 - 375 of 1024 in total

Filters

Subject Area

Life Sciences (627)

Physical Sciences (148)

Social Sciences (148)

Technology and Engineering (87)

Uncategorized

Arts and Humanities (2)

Funder

Other (275)

U.S. Department of Energy (DOE) (249)

U.S. National Science Foundation (NSF) (244)

U.S. National Institutes of Health (NIH) (89)

U.S. Department of Agriculture (USDA) (64)

Illinois Department of Natural Resources (IDNR) (26)

U.S. Geological Survey (USGS) (8)

U.S. National Aeronautics and Space Administration (NASA) (6)

Illinois Department of Transportation (IDOT) (4)

U.S. Army (3)

Publication Year

2025 (286)

2021 (108)

2022 (106)

2024 (105)

2020 (96)

2023 (75)

2019 (72)

2018 (61)

2026 (44)

2017 (36)

2016 (30)

2009 (1)

2011 (1)

2012 (1)

2014 (1)

2015 (1)

License

CC BY (528)

CC0 (469)

custom (27)

Illinois Data Bank Dataset Search Results

Results

published: 2024-05-23

Data for: Learned 1-D passive scalar advection to accelerate chemical transport modeling: a case study with GEOS-FP horizontal wind fields

Park, Manho; Zheng, Zhonghua; Riemer, Nicole; Tessum, Christopher (2024)

This dataset contains the training results (model parameters, outputs), datasets for generalization testing, and 2-D implementation used in the article "Learned 1-D passive scalar advection to accelerate chemical transport modeling: a case study with GEOS-FP horizontal wind fields." The article will be submitted to Artificial Intelligence for Earth Systems. The datasets are saved as CSV for 1-D time-series data and *netCDF for 2-D time series dataset. The model parameters are saved in every training epoch tested in the study.

keywords: Air quality modeling; Coarse-graining; GEOS-Chem; Numerical advection; Physics-informed machine learning; Transport operator

published: 2025-09-30

Data from Economic Perspective of Ethanol and Biodiesel Coproduction from Industrial Hemp

Viswanathan, Mothi Bharath; Cheng, Ming-Hsun; Clemente, Tom; Dweikat, Ismail; Singh, Vijay (2025)

In this study, the economics of producing biofuels from an industrial hemp (Cannabis sativa) genotype – 19m96136 was investigated. A lignocellulosic biofuel plant, hourly consuming 85 metric tons of hemp biomass was modeled in SuperPro Designer®. The integrated bioenergy plant produced hemp biodiesel and bioethanol from lipids and carbohydrates, respectively. The structural composition of the industrial hemp plant was analyzed in a previous study. The data obtained was used to simulate feedstock composition in SuperPro Designer®. The simulation results indicated that Hemp containing 2% lipids can yield up to 3.95 million gallons of biodiesel annually. On improving biomass lipid content to 5 and 10%, biodiesel production increased to 9.88 and 19.91 million gallons, respectively. The breakeven unit production cost of hemp biodiesel with 2, 5, and 10% lipid containing hemp was $18.49, $7.87, and $4.13/gallon, respectively. The biodiesel unit production cost when utilizing 10% lipid-containing hemp was comparable to soybean biodiesel at $4.13/gallon. Furthermore, sensitivity analysis revealed the possibility of a 7.80% reduction in unit production cost upon a 10% reduction in hemp feedstock cost. Furthermore, industrial hemp was capable of producing between 307.80 and 325.82 gallons of total biofuels per hectare of agricultural land than soybean.

keywords: Conversion;Feedstock Production;Economics;Modeling

published: 2018-06-06

Dataset for Evaluation of DeNitrification DeComposition Model for Estimating Ammonia Fluxes from Chemical Fertilizer Application

Balasubramanian, Srinidhi; Nelson, Andrew; Koloutsou-Vakakis, Sotiria; Lin, Jie; Rood, Mark; Myles, LaToya; Bernacchi, Carl (2018)

DNDC scripts and outputs that were generated as a part of the research publication 'Evaluation of DeNitrification DeComposition Model for Estimating Ammonia Fluxes from Chemical Fertilizer Application'.

keywords: DNDC; REA; ammonia emissions; fertilizers; uncertainty analysis

published: 2024-03-25

Data for "Differing physiological performance of coexisting cool- and warmwater fish species under heatwaves in the Midwestern United States"

Suski, Cory; Dai, Qihong (2024)

This is the dataset for the manuscript titled, "Differing physiological performance of coexisting cool- and warmwater fish species under heatwaves in the Midwestern United States"

keywords: climate change; heat wave; metabolic rate; swimming; predator-prey interaction; thermal tolerance; Sander vitreus; walleye; largemouth bass; species distributions

published: 2024-06-24

Data for “Autophagy suppression in DNA damaged cells occurs through a newly identified p53-proteasome-LC3 axis”

Lieu, D'Feau J.; Crowder, Molly K.; Kryza, Jordan R.; Tamilselvam, Batcha; Kaminski, Paul J.; Kim, Ik-Jung; Li, Yingxing; Jeong, Eunji; Enkhbaatar, Michidmaa; Chen, Henry; Son, Sophia B.; Mok, Hanlin; Bradley, Kenneth A.; Phillips, Heidi; Blanke, Steven R. (2024)

This page contains the data for the manuscript "Autophagy suppression in DNA damaged cells occurs through a newly identified p53-proteasome-LC3 axis" currently available in preprint on bioRxiv

keywords: Steven R Blanke; Cytolethal Distending Toxin; CDT; Autophagy; Genotoxicity; p53; DNA damage; DNA damage response; LC3; proteasome; proteostasis; DDR; autophagosome

published: 2024-07-09

Data matrices for "Missing Data and Model Selection in Phylogenomics: A Re-Evaluation of Cicadomorpha (Hemiptera: Auchenorrhyncha) Superfamily Level Relationships Under Site-Heterogeneous Models"

Yan, Bin; Dietrich, Christopher; Yu, Xiaofei; Jiang, Yan; Dai, Renhuai; Du, Shiyu; Cai, Chenyang; Yang, Maofa; Zhang, Feng (2024)

The included files are the alignments of DNA or amino acid sequences used for phylogenetic analyses of Auchenorrhyncha (Insecta: Hemiptera) in the manuscript by Bin et al. submitted to the journal “Systematic Entomology.” The files are plain text in either FASTA (.fa or .fas suffix) or PHYLIP (.phy suffix) format. Matrix0 is the set of all loci after multiple sequence alignment and trimming (hereafter called). Matrix1 consists of loci having 75% average bootstrap support and 80% taxon completeness (hereafter called Matrix1). Matrix2 consists of loci having 75% average bootstrap support and 95% completeness. Matrix2_nt12 is the same as Matrix2 but with third codon positions excluded. More details on how the datasets were compiled is provided in the Methods section of the manuscript file, also included as a PDF. Supplemental figures for the submitted manuscript are also provided as a PDF for additional information.

keywords: Insecta; Phylogeny; DNA sequence; Evolution

published: 2018-03-08

Molecular Biology Databases Published in Nucleic Acids Research between 1991-2016

Imker, Heidi (2018)

This dataset was developed to create a census of sufficiently documented molecular biology databases to answer several preliminary research questions. Articles published in the annual Nucleic Acids Research (NAR) “Database Issues” were used to identify a population of databases for study. Namely, the questions addressed herein include: 1) what is the historical rate of database proliferation versus rate of database attrition?, 2) to what extent do citations indicate persistence?, and 3) are databases under active maintenance and does evidence of maintenance likewise correlate to citation? An overarching goal of this study is to provide the ability to identify subsets of databases for further analysis, both as presented within this study and through subsequent use of this openly released dataset.

keywords: databases; research infrastructure; sustainability; data sharing; molecular biology; bioinformatics; bibliometrics

published: 2025-06-23

Data for ESMDynamic: Fast and Accurate Prediction of Protein Dynamic Contact Maps from Single Sequences

Kleiman, Diego; Feng, Jiangyan; Xue, Zhengyuan; Shukla, Diwakar (2025)

This repository contains data and model weights associated with the publication "ESMDynamic: Fast and Accurate Prediction of Protein Dynamic Contact Maps from Single Sequences". It includes the datasets used for training and evaluating a dynamic contact prediction model, ESMDynamic, as well as a script for conversion and usage.

keywords: Computational biology; Structural biology; Molecular dynamics; Machine learning; Protein modeling; Bioinformatics; Biophysics; Artificial intelligence

published: 2025-05-05

Data for article about perceived trustworthiness of social media AI-Generated content

Benson, Sara; Cheng, Siyao; Ton, Mary; Graves, Celenia; Owens, Dawn (2025)

The dataset includes responses from approximately 550 participants to survey questions about trust in images labeled with AI-related tags, compared to other images found online. The questions also explore how the type of label influences their trust.

keywords: Artificial intelligence (AI); Trust in AI; Al labeling; AI ethics

published: 2016-06-06

Datasets for modeling collaborative formation and collaborative "success"

Fegley, Brent D. (2016)

These datasets represent first-time collaborations between first and last authors (with mutually exclusive publication histories) on papers with 2 to 5 authors in years [1988,2009] in PubMed. Each record of each dataset captures aspects of the similarity, nearness, and complementarity between two authors about the paper marking the formation of their collaboration.

published: 2016-12-13

Sequencing data for motility selection experiments

Fraebel, David T.; Kuehn, Seppe (2016)

BAM files for founding strain (MG1655-motile) as well as evolved strains from replicate motility selection experiments in low-viscosity agar plates containing either rich medium (LB) or minimal medium (M63+0.18mM galactose)

published: 2019-05-22

Isolated artificial spin ice kinetics

Lao, Yuyang; Schiffer, Peter (2019)

This is the experimental data of isolated nanomagnet islands with or without the presence of large nanomagnet islands. The small islands are made of Permalloy materials with size of 170 nm by 470 nm by 2.5 nm. The systems are measured at a temperature where the small islands are fluctuating around room temperature. The data is recorded as photoemission electron microscopy intensity. More details about the data can be found in the note.txt and Spe_2016.xlsx file. Note: The raw data folders are stored in five volumes during the compression. All five volumes are needed in order to recover the original folder.

keywords: artificial spin ice; magnetism

published: 2020-12-29

Fern functional traits

Viana, Jéssica; Turner, Benjamin; Dalling, James (2020)

Three datasets: species_abundance_data, species_traits, and environmental_data. The three datasets were collected in the Fortuna Forest Reserve (8°45′ N, 82°15′ W) and Palo Seco Protected Forest (8°45′ N, 82°13′ W) located in western Panama. The two reserves support humid to super-humid rainforests, according to Holdridge (1947). The species_abundance_data and species_traits datasets were collected across 15 subplots of 25 m2 in 12 one-hectare permanent plots distributed across the two reserves. The subplots were spaced 20 m apart along three 5 m wide transects, each 30 m apart. Please read Prada et al. (2017) for details on the environmental characteristics of the study area. Prada CM, Morris A, Andersen KM, et al (2017) Soils and rainfall drive landscape-scale changes in the diversity and functional composition of tree communities in a premontane tropical forest. J Veg Sci 28:859–870. https://doi.org/10.1111/jvs.12540

keywords: functional traits; plants; ferns; environmental data; Fortuna; species data; community ecology

published: 2021-09-03

Dataset for evaluating the Hind/He statistic in polyRAD

Clark, Lindsay V.; Mays, Wittney; Lipka, Alexander E.; Sacks, Erik J. (2021)

All of the files in this dataset pertain to the evaluation of a novel statistic, Hind/He, for distinguishing Mendelian loci from paralogs. They are derived from a RAD-seq genotyping dataset of diploid and tetraploid Miscanthus sacchariflorus.

published: 2021-08-20

Maize and Sorghum Establishment and Yield following Pre-Emergence Waterlogging

von Haden, Adam C.; DeLucia, Evan H.; Yang, Wendy; Burnham, Mark (2021)

In 2020, early-season extreme precipitation events occurred following the planting of Sorghum bicolor (L.) Moench and Zea mays L. in central Illinois that caused ponding. Following the first rainfall event 50m transects were established to assess the waterlogging effects on seedling emergence and crop yields. Soil moisture, emergence, stem and tiller count, LAI, and yield were measured at various points in the season along these transects.

keywords: Sorghum; Maize; Emergence; Yield; LAI

published: 2022-02-11

FASTA file of the final sequence alignment used in the haplotype analyses of Culex pipiens complex populations collected in south-eastern Illinois (2016-2017)

Trivellone, Valeria; Cao, Yanghui; Blackshear, Millon; Kim, Chang-Hyun; Stone, Christopher (2022)

The Culex_Trivellone_etal.fas fasta file contains the original final sequence alignment used in the haplotype analyses of Trivellone et al. (Frontiers in Public Health, under review). The 492 sequences (from specimens of Culex pipiens complex collected in different habitat types using a BG-sentinel traps) were aligned using PASTA v1.8.5 under default settings. The final dataset contains 686 positions of the cytochrome c oxidase subunit I (COI) mitochondrial gene. The data analyses are further described in the cited original paper.

keywords: Culex; Culicidae; COI; mosquito surveillance, species assemblages

published: 2024-03-27

Dataset for "Arguing about Controversial Science in the News: Does Epistemic Uncertainty Contribute to Information Disorder?"

Zheng, Heng; Schneider, Jodi (2024)

To gather news articles from the web that discuss the Cochrane Review, we used Altmetric Explorer from Altmetric.com and retrieved articles on August 1, 2023. We selected all articles that were written in English, published in the United States, and had a publication date <b>prior to March 10, 2023</b> (according to the “Mention Date” on Altmetric.com). This date is significant as it is when Cochrane issued a statement about the "misleading interpretation" of the Cochrane Review. The collection of news articles is presented in the Altmetric_data.csv file. The dataset contains the following data that we exported from Altmetric Explorer: - Publication date of the news article - Title of the news article - Source/publication venue of the news article - URL - Country We manually checked and added the following information: - Whether the article still exists - Whether the article is accessible - Whether the article is from the original source We assigned MAXQDA IDs to the news articles. News articles were assigned the same ID when they were (a) identical or (b) in the case of Article 207, closely paraphrased, paragraph by paragraph. Inaccessible items were assigned a MAXQDA ID based on their "Mention Title". For each article from Altmetric.com, we first tried to use the Web Collector for MAXQDA to download the article from the website and imported it into MAXQDA (version 22.7.0). If an article could not be retrieved using the Web Collector, we either downloaded the .html file or in the case of Article 128, retrieved it from the NewsBank database through the University of Illinois Library. We then manually extracted direct quotations from the articles using MAXQDA. We included surrounding words and sentences, and in one case, a news agency’s commentary, around direct quotations for context where needed. The quotations (with context) are the positions in our analysis. We also identified who was quoted. We excluded quotations when we could not identify who or what was being quoted. We annotated quotations with codes representing groups (government agencies, other organizations, and research publications) and individuals (authors of the Cochrane Review, government agency representatives, journalists, and other experts such as epidemiologists). The MAXQDA_data.csv file contains excerpts from the news articles that contain the direct quotations we identified. For each excerpt, we included the following information: - MAXQDA ID of the document from which the excerpt originates; - The collection date and source of the document; - The code with which the excerpt is annotated; - The code category; - The excerpt itself.

keywords: altmetrics; MAXQDA; polylogue analysis; masks for COVID-19; scientific controversies; news articles

published: 2021-11-18

Rewritable Two-Dimensional DNA-Based Data Storage System (2DDNA) Sequencing Dataset

Pan, Chao; Tabatabaei, S Kasra; Tabatabaei Yazdi, S. M. Hossein; Hernandez, Alvaro; Schroeder, Charles; Milenkovic, Olgica (2021)

This dataset contains sequencing data obtained from Illumina MiSeq device to prove the concept of the proposed 2DDNA framework. Please refer to README.txt for detailed description of each file.

keywords: machine learning;image processing;computer vision;rewritable storage system;2D DNA-based data storage

published: 2025-01-31

Airyscan confocal superresolution images of extant Malvaceae pollen with a focus on Bombacoideae

Punyasena, Surangi W.; Romero, Ingrid; Urban, Michael A. (2025)

Title: Airyscan confocal superresolution images of extant Malvaceae pollen with a focus on Bombacoideae Authors: Surangi W. Punyasena, Ingrid Romero, Michael A. Urban Subject: Biological sciences Keywords: Malvaceae; superresolution microscopy; Zeiss; Bombacacidites; Neotropics; CZI Funder: NSF-DBI Advances in Bioinformatics (NSF-DBI-1262561) Corresponding Creator: Surangi W. Punyasena This dataset includes a total of 430 images of extant specimens of the Malvaceae, with a focus on species that are or have been included within the subfamily Bombacoideae. There are 27 genera included within 26 folders. Each folder is named by genus and contains all the images that correspond to that genus. Note that the genus _Matisia_ is included with _Quararibea_ as detailed in the metadata READ ME file. The specimens imaged are from the palynological collections of the Swedish Museum of Natural History and Smithsonian Tropical Research Institute, and herbarium specimens from the Smithsonian Herbarium National Museum. The optical superresolution microscopy images were taken using a Zeiss LSM 880 with Airyscan at 630X magnification (63x/NA 1.4 oil DIC). The images are in the original CZI file format. They can be opened using Zeiss propriety software (Zen, Zen lite) or in ImageJ/FIJI. More information on how to open CZI files can be found here: [https://www.zeiss.com/microscopy/en/products/software/zeiss-zen/czi-image-file-format.html] Image metadata and file organization are described in the CSV file "METADATA_Malvaceae_Bombacoideae_modern-species.csv". The column headings are: Folder The folder in which the image file is found Subfamily The current subfamily determination based on the literature. Note that _Pentaplaris_ and _Septotheca_ have not been assigned a subfamily. Genus Genus name Species Species name Accepted name Accepted species name, updated from the literature Slide name Species name as denoted on the herbarium slide Collection Source of the herbarium slide: Sweden National Museum of Natural History or the Smithsonian Tropical Research Institute File name File name using the species name denoted on the herbarium slide Slide ID/Herbarium ID Specimen collection number Please cite this dataset as: Punyasena, Surangi W.; Romero, Ingrid; Urban, Michael A. (2025): Airyscan confocal superresolution images of extant Malvaceae pollen with a focus on Bombacoideae. University of Illinois Urbana-Champaign. https://doi.org/10.13012/B2IDB-2968712_V1

keywords: Malvaceae; superresolution microscopy; Zeiss; Bombacoideae; Neotropics; CZI

published: 2015-12-16

Data for Ultra-Large Alignments Using Phylogeny-Aware Profiles

Nguyen, Nam-phuong; Mirarab, Siavash; Kumar, Keerthana; Warnow, Tandy (2015)

This dataset contains the data for PASTA and UPP. PASTA data was used in the following articles: Mirarab, Siavash, Nam Nguyen, Sheng Guo, Li-San Wang, Junhyong Kim, and Tandy Warnow. “PASTA: Ultra-Large Multiple Sequence Alignment for Nucleotide and Amino-Acid Sequences.” Journal of Computational Biology 22, no. 5 (2015): 377–86. doi:10.1089/cmb.2014.0156. Mirarab, Siavash, Nam Nguyen, and Tandy Warnow. “PASTA: Ultra-Large Multiple Sequence Alignment.” Edited by Roded Sharan. Research in Computational Molecular Biology, 2014, 177–91. UPP data was used in: Nguyen, Nam-phuong D., Siavash Mirarab, Keerthana Kumar, and Tandy Warnow. “Ultra-Large Alignments Using Phylogeny-Aware Profiles.” Genome Biology 16, no. 1 (December 16, 2015): 124. doi:10.1186/s13059-015-0688-z.

published: 2025-10-30

Data for Construction of a Compact Array of Microplasma Jet Devices and Its Application for Random Mutagenesis of Rhodosporidium toruloides

Koh, Hyun Gi; Kim, Jinhong; Rao, Christopher V.; Park, Sung-Jin; Jin, Yong-Su (2025)

A small and efficient DNA mutation-inducing machine was constructed with an array of microplasma jet devices (7 × 1) that can be operated at atmospheric pressure for microbial mutagenesis. Using this machine, we report disruption of a plasmid DNA and generation of mutants of an oleaginous yeast Rhodosporidium toruloides. Specifically, a compact-sized microplasma channel (25 × 20 × 2 mm3) capable of generating an electron density of greater than 1013 cm–3 was constructed to produce reactive species (N2*, N2+, O, OH, and Hα) under helium atmospheric conditions to induce DNA mutagenesis. The length of microplasma channels in the device played a critical role in augmenting both the volume of plasma and the concentration of reactive species. First, we confirmed that microplasma treatment can linearize a plasmid by creating nicks in vitro. Second, we treated R. toruloides cells with a jet device containing 7 microchannels for 5 min; 94.8% of the treated cells were killed, and 0.44% of surviving cells showed different colony colors as compared to their parental colony. Microplasma-based DNA mutation is energy-efficient and can be a safe alternative for inducing mutations compared to conventional methods using toxic mutagens. This compact and scalable device is amenable for industrial strain improvement involving large-scale mutagenesis.

keywords: Conversion;Genome Engineering

published: 2020-02-12

Data for: Auditing Race and Gender Discrimination in Online Housing Markets

Asplund, Joshua; Karahalios, Karrie (2020)

This dataset contains the results of a three month audit of housing advertisements. It accompanies the 2020 ICWSM paper "Auditing Race and Gender Discrimination in Online Housing Markets". It covers data collected between Dec 7, 2018 and March 19, 2019. There are two json files in the dataset: The first contains a list of json objects representing advertisements separated by newlines. Each object includes the date and time it was collected, the image and title (if collected) of the ad, the page on which it was displayed, and the training treatment it received. The second file is a list of json objects representing a visit to a housing lister separated by newlines. Each object contains the url, training treatment applied, the location searched, and the metadata of the top sites scraped. This metadata includes location, price, and number of rooms. The dataset also includes the raw images of ads collected in order to code them by interest and targeting. These were captured by selenium and named using a perceptive hash to de-duplicate images.

keywords: algorithmic audit; advertisement audit;

published: 2020-05-31

Simulated multi-sample tumor bulk sequencing data

Zhang, Chuanyi; El-Kebir, Mohammed; Ochoa, Idoia (2020)

This repository includes a simulated dataset and related scripts used for the paper "Moss: Accurate Single-Nucleotide Variant Calling from Multiple Bulk DNA Tumor Samples".

keywords: Somatic Mutations; Bulk DNA Sequencing; Cancer Genomics

published: 2022-02-14

Data for: Quantifying the effects of mixing state on aerosol optical properties

Yao, Yu; Curtis, Jeffrey; Ching, Joseph; Zheng, Zhonghua; Riemer, Nicole (2022)

This dataset contains simulation results from numerical model PartMC-MOSAIC used in the article "Quantifying the effects of mixing state on aerosol optical properties". This article is submitted to the journal Atmospheric Physics and Chemistry. There are total 100 scenario directories in this dataset, denoted from 00-99. Each scenario contains 25 NetCDF files hourly output from PartMC-MOSAIC simulations containing the simulated gas and particle information. The data was produced using version 2.5.0 of PartMC-MOSAIC. Instructions to compile and run PartMC-MOSAIC are available at https://github.com/compdyn/partmc. The chemistry code MOSAIC is available by request from Rahul.Zaveri@pnl.gov. For more details of reproducing the cases, please contact nriemer@illinois.edu and yuyao3@illinois.edu.

keywords: Aerosol mixing state; Aerosol optical properties; Mie calculation; Black Carbon

published: 2025-06-22

Microclimate Species Distribution Models and Mechanistic Models of Potential Surface Activity for Plethodontid Salamanders in Great Smoky Mountains National Park

Stickley, Samuel; Crawford, John; Peterman, William; Fraterrigo, Jennifer (2025)

keywords: terrestrial salamanders, microhabitat, physiology, mechanistic models, ecological niche models, climate change, Great Smoky Mountains National Park