Dataset Search

Displaying 76 - 100 of 473 in total

Filters

Subject Area

Life Sciences (282)

Social Sciences (84)

Physical Sciences (67)

Technology and Engineering (37)

Uncategorized

Arts and Humanities (1)

Funder

U.S. Department of Energy (DOE) (150)

Other (116)

U.S. National Science Foundation (NSF) (112)

U.S. National Institutes of Health (NIH) (37)

U.S. Department of Agriculture (USDA) (28)

Illinois Department of Natural Resources (IDNR) (12)

U.S. Geological Survey (USGS) (2)

Illinois Department of Transportation (IDOT) (1)

U.S. National Aeronautics and Space Administration (NASA) (1)

U.S. Army (1)

Publication Year

2025 (153)

2022 (50)

2024 (50)

2021 (45)

2020 (36)

2023 (34)

2018 (29)

2026 (28)

2019 (27)

2016 (11)

2017 (10)

License

CC BY (267)

CC0 (194)

custom (12)

Illinois Data Bank Dataset Search Results

Results

published: 2018-07-29

NJMerge: A generic technique for scaling phylogeny estimation methods and its application to species trees

Molloy, Erin K.; Warnow, Tandy (2018)

This repository includes scripts, datasets, and supplementary materials for the study, "NJMerge: A generic technique for scaling phylogeny estimation methods and its application to species trees", presented at RECOMB-CG 2018. The supplementary figures and tables referenced in the main paper can be found in njmerge-supplementary-materials.pdf. The latest version of NJMerge can be downloaded from Github: https://github.com/ekmolloy/njmerge. ***When downloading datasets, please note that the following errors.*** In README.txt, lines 37 and 38 should read: + fasttree-exon.tre contains lines 1-25, 1-100, or 1-1000 of fasttree-total.tre + fasttree-intron.tre contains lines 26-50, 101-200, or 1001-2000 of fasttree-total.tre Note that the file names (fasttree-exon.tre and fasttree-intron.tre) are swapped. In tools.zip, the compare_trees.py and the compare_tree_lists.py scripts incorrectly refer to the "symmetric difference error rate" as the "Robinson-Foulds error rate". Because the normalized symmetric difference and the normalized Robinson-Foulds distance are equal for binary trees, this does not impact the species tree error rates reported in the study. This could impact the gene tree error rates reported in the study (see data-gene-trees.csv in data.zip), as FastTree-2 returns trees with polytomies whenever 3 or more sequences in the input alignment are identical. Note that the normalized symmetric difference is always greater than or equal to the normalized Robinson-Foulds distance, so the gene tree error rates reported in the study are more conservative. In njmerge-supplementary-materials.pdf, the alpha parameter shown in Supplementary Table S2 is actually the divisor D, which is used to compute alpha for each gene as follows. 1. For each gene, a random value X between 0 and 1 is drawn from a uniform distribution. 2. Alpha is computed as -log(X) / D, where D is 4.2 for exons, 1.0 for UCEs, and 0.4 for introns (as stated in Table S2). Note that because the mean of the uniform distribution (between 0 and 1) is 0.5, the mean alpha value is -log(0.5) / 4.2 = 0.16 for exons, -log(0.5) / 1.0 = 0.69 for UCEs, and -log(0.5) / 0.4 = 1.73 for introns.

keywords: phylogenomics; species trees; incomplete lineage sorting; divide-and-conquer

published: 2025-11-19

Data for Redefining the Product Portfolio of Oilcane Bagasse Biorefinery: Recovering Natural Colorants, Vegetative Lipids and Sugars

Banerjee, Shivali; Beraja, Galit; Eilts, Kristen; Singh, Vijay (2025)

:Bioenergy crops have been known for their ability to produce biofuels and bioproducts. In this study, the product portfolio of recently developed transgenic sugarcane (oilcane) bagasse has been redefined for recovering natural pigments (anthocyanins), sugars, and vegetative lipids. The total anthocyanin content in oilcane bagasse has been estimated as 92.9 ± 18.9 µg/g of dried bagasse with cyanidin-3-glucoside (13.5 ± 18.9 µg per g of dried bagasse) as the most prominent anthocyanin present. More than 85 % (w/w) of the total anthocyanins were recovered from oilcane bagasse at a pretreatment temperature of 150 °C for 15 min. These conditions for the hydrothermal pretreatment also led to a 2-fold increase in the glucose yield upon the enzymatic saccharification of the pretreated bagasse. Further, a 1.5-fold enrichment of the vegetative lipids was demonstrated in the pretreated residue. Re-defining green biorefineries with multiple high-value products in a zero-waste approach is the need of the hour for attaining sustainability.

keywords: Conversion;Biomass Analytics;Bioproducts;Biorefinery;Oilcane

published: 2022-04-15

Data on "Evaluation of CO2 sealing potential of heterogeneous Eau Claire shale"

Kim, Hyunbin; Makhnenko, Roman (2022)

This dataset is provided to support the statements in Kim, H., and R.Y. Makhnenko. 2022. "Evaluation of CO2 sealing potential of heterogeneous Eau Claire shale". Journal of the Geological Society. In geologic carbon dioxide (CO2) storage in deep saline aquifers, buoyant CO2 tends to float upwards in the reservoirs overlaid by low permeable formations called caprocks. Caprocks should serve as barriers to potential CO2 leakage that can happen through a diffusion loss and permeation through faults, fractures, or pore spaces. The leakage through intact caprock would mainly depend on its permeability and CO2 breakthrough pressure, and is affected by the heterogeneities in the material. Here, we study the sealing potential of a caprock from Illinois Basin - Eau Claire shale, with sandy and shaly fractions distinguished via electron microscopy and grain/pore size and surface area characterization. The direct measurements of permeability of sandy shale provides the values ~ 10-15 m2, while clayey specimens are three orders of magnitude less permeable. The CO2 breakthrough pressure under in-situ stress conditions is 0.1 MPa for the sandy shale and 0.4 MPa for the clayey counterpart – these values are higher than those predicted by the porosimetry methods performed on the unconfined specimens. Sandy Eau Claire shale would allow penetration of large CO2 volumes at low overpressures, while the clayey formation can serve as a caprock in the absence of faults and fractures in it.

keywords: Geologic carbon storage; Caprock; Shale; CO2 breakthrough pressure; Porosimetry.

published: 2025-11-12

Data for Rapid and Efficient in planta Genome Editing in Sorghum Using Foxtail Mosaic Virus-mediated sgRNA Delivery

BAYSAL, CAN; Kausch, Albert P.; Cody, Jon P.; Altpeter, Fredy; Voytas, Daniel (2025)

The requirement of in vitro tissue culture for the delivery of gene editing reagents limits the application of gene editing to commercially relevant varieties of many crop species. To overcome this bottleneck, plant RNA viruses have been deployed as versatile tools for in planta delivery of recombinant RNA. Viral delivery of single-guide RNAs (sgRNAs) to transgenic plants that stably express CRISPR-associated (Cas) endonuclease has been successfully used for targeted mutagenesis in several dicotyledonous and few monocotyledonous plants. Progress with this approach in monocotyledonous plants is limited so far by the availability of effective viral vectors. We engineered a set of foxtail mosaic virus (FoMV) and barley stripe mosaic virus (BSMV) vectors to deliver the fluorescent protein AmCyan to track viral infection and movement in Sorghum bicolor. We further used these viruses to deliver and express sgRNAs to Cas9 and Green Fluorescent Protein (GFP) expressing transgenic sorghum lines, targeting Phytoene desaturase (PDS), Magnesium-chelatase subunit I (MgCh), 4-hydroxy-3-methylbut-2-enyl diphosphate reductase, orthologs of maize Lemon white1 (Lw1) or GFP. The recombinant BSMV did neither infect sorghum nor deliver or express AmCyan and sgRNAs. In contrast, the recombinant FoMV systemically spread throughout sorghum plants and induced somatic mutations with frequencies reaching up to 60%. This mutagenesis led to visible phenotypic changes, demonstrating the potential of FoMV for in planta gene editing and functional genomics studies in sorghum.

keywords: Feedstock Production;Genome Engineering;Genomics

published: 2024-04-11

The missing phosphorus legacy of the Anthropocene: quantifying residual phosphorus in the biosphere

Margenot, Andrew; Zhou, Shengnan; Xu, Suwei; Condron, Leo; Metson, Geneviève; Haygarth, Philip; Wade, Jordon; Agyeman, Price Chapman (2024)

A defining feature of the Anthropocene is the distortion of the biosphere phosphorus (P) cycle. A relatively sudden acceleration of input fluxes without a concomitant increase in output fluxes has led to net accumulation of P in the terrestrial-aquatic continuum. Over the past century, P has been mined from geological deposits to produce crop fertilizers. When P inputs are not fully removed with harvest of crop biomass, the remaining P accumulates in soils. This residual P is a uniquely anthropogenic pool of P, and its management is critical for agronomic and environmental sustainability. This dataset includes data for us to quantify residual P from different long-term managed systems. The following is the desccription of the dataset. There are 7 sheets in total. 1. P_balance: From Morrow Plots maize-maize rotaiton (1888-2021), L: Low estimation; M: medium estimation; H: high estimation; 2. M3P: From Morrow Plots selected plots (selected years), M3P_sur: Mehlich III P concentration in surface 17cm soils; M3P_sub: Mehlich III P concentration in 17-34cm subsoils; P_balance: the difference between P inputs and P outputs; TP_sur: total P stocks in surface 17cm soils; TP_sub: total P stocks in 17-34cm subsoils; 3. Morrow_Plot_P_pool_all: Group: a - labile P; b - Fe/Al-P; c - Ca-P; d - total organic P; e - non-extractable P; Fertilized: P stocks in the fertilized plot; Unfertilized: P stocks in the unfertilized plot; F-U: difference between P stocks in ther fertilized and unfertilized plots; dif%: percent difference in total P; 4. Rothamsted_P_pool_all: Treatment: Unfertilized: no fertilization; FYM: farmyard manure; PK: synthetic P and K fertilizer; Group: a - labile P; b - Fe/Al-P; c - Ca-P; d - total organic P; e - non-extractable P; P_change: differnce in P stocks over time; dif%: percent difference in total P; 5. L'Acadie_P_pool_all: Treatment: MP_LowP: moldboard plow with low rate of P fertilizer; MP_HighP: moldboard plow with high rate of P fertilizer; NT_LowP: no till with low rate of P fertilizer; NT_HighP: no till with high rate of P fertilizer; Group: a - labile P; b - Fe/Al-P; c - Ca-P; d - total organic P; e - non-extractable P; P_change: differnce in P stocks over time; dif%: percent difference in total P; 6. Rothamsted_P_pool_duration: Treatment: Unfertilized: no fertilization; FYM: farmyard manure; PK: synthetic P and K fertilizer; Duration: from a year to another year; Group: a - labile P; b - Fe/Al-P; c - Ca-P; d - total organic P; e - non-extractable P; P_change: differnce in P stocks over time; dif%: percent difference in total P; 7. L'Acadie_P_pool_duration: Treatment: MP_LowP: moldboard plow with low rate of P fertilizer; MP_HighP: moldboard plow with high rate of P fertilizer; NT_LowP: no till with low rate of P fertilizer; NT_HighP: no till with high rate of P fertilizer; Duration: from a year to another year; Group: a - labile P; b - Fe/Al-P; c - Ca-P; d - total organic P; e - non-extractable P; P_change: differnce in P stocks over time; dif%: percent difference in total P;

keywords: phosphate rock; biosphere; balances; soil test P; long-term experiment

published: 2025-12-02

Data for The Effects of Sequential Hydrothermal-Mechanical Refining Pretreatment on Cellulose Structure Changes and Sugar Recoveries

Cheng, Ming-Hsun; Maitra, Shraddha; Carr Clennon, Aidan N.; Appell, Michael; Dien, Bruce; Singh, Vijay (2025)

The recalcitrance of lignocellulosic biomass necessitates an efficient pretreatment protocol for operating a successful cellulosic biorefinery. It is critical to improve cellulose accessibility for hydrolysis and fermentation by altering the plant cell wall’s physical structure and chemical composition. Sequential hydrothermal-mechanical refining pretreatment (HMR) allows efficient recovery of cellulosic sugars without utilizing any hazardous chemicals. HMR has been successfully applied to Liberty switchgrass, a bioenergy cultivar released by the USDA, and now it is being applied to oilcane, a recently developed transgenic sugarcane variety engineered to accumulate lipids in its vegetative tissues. Sugar yields of oilcane bagasse (OCB) and switchgrass (SG) treated with HMR are 96.4% and 75.4%, respectively. This study sought to correlate cellulosic sugar yields with structural changes within the cell wall caused by HMR on two distinct bioenergy crops. Simon’s staining technique for the specific surface area analysis showed that HMR increased the specific surface area of pretreated biomass residues by 80-112%. In addition, ATR-FTIR was performed to determine the effects of HMR on physical structures based on the total crystallinity index (TCI) and hydrogen bonding intensity (HBI). Irrespective of biomass type, HMR decreased the initial crystalline cellulose contents of untreated biomass residues by 3.5% and reduced TCI and HBI by 7-13%. The study found that sugar yields were negatively correlated to reducing values of hydrogen bonding intensity, crystalline cellulose content, and total crystallinity index.

keywords: Conversion;Biomass Analytics;Economics;Hydrolysate

published: 2019-03-19

Meltwater Meandering Channels on Ice: Centerlines and Images

Fernandez, Roberto; Parker, Gary; Stark, Colin P. (2019)

This dataset includes images and extracted centerlines from experiments looking at the formation and evolution of meltwater meandering channels on ice. The laboratory data includes centimeter- and millimeter-scale rivulets. Dataset also includes an image and corresponding centerlines from the Peterman Ice Island. All centerlines were manually digitized in Matlab but no distributable code was developed for the process. Once digitized, centerlines were smoothed and standardized following methods and routines developed by other authors (Zolezzi and Guneralp, 2016; Guneralp and Rhoads, 2008). Details about the preparation of the centerlines and processing with these methods is included in the dissertation by Fernández (2018) linked to this dataset. "Millimeter scale and Peterman Ice Island centerlines.pdf": This file includes the images of two mm-scale experimetns and the Peterman Ice Island image. Seventeen centerlines were digitized from the former and seven were digitized from the latter. Those centerlines are shown above the images themselves. "Centimeter scale rivulet images.pdf": This file includes images corresponding to all cm-scale centerlines used for the analysis presented in the dissertation by Fernandez (2018). Each image has a short caption indicating the run ID and the time at which it was captured. The images were used to extract centerlines to look at the planform evolution of cm-scale meltwater meandering rivulets on ice. Images include 26 centerlines from four different runs. "Meltwater meandering channel centerlines.xlsx": This spreadsheet contains the centerline data for all fifty centerlines. The workbook includes 51 sheets. The first 50 are related to each one of the channels. The mm scale and Peterman Ice Island ones are identified using the same IDs shown in "Millimeter scale and Peterman Ice Island centerlines.pdf". The cm-scale centerlines are identified by run ID and a number indicating the time in minutes (with t = 0 min being the time at which water started flowing over the ice block). The naming convention is also associated to the images in "Centimeter scale rivulet images.pdf". The last sheet in the workbook includes a summary of the channel widths measured from every image for each centerline. The 50 sheets with the centerline information have four columns each. The titles of the columns are X, Y, S, and C. X,Y are dimensionless coordinates of the centerline. S is dimensionless streamwise coordinate (location along the centerline). C is dimensionless curvature value. All these values were non-dimensionalized with the channel width. See Fernandez (2018), Zolezzi and Guneralp (2016), and Guneralp and Rhoads (2008) for more details regarding the process of smoothing, standardizing and non-dimensionalization of the centerline coordinates.

keywords: Meltwater, Meandering, Ice, Supraglacial, Experiments

published: 2025-08-20

Comparative economic analysis between bioenergy and forage types of switchgrass for sustainable biofuel feedstock production: A DEA and cost-benefit analysis approach

Arshad, Muhammad Umer; Archer, David ; Wasonga, Daniel ; Namoi, Nictor; Boe, Arvid ; Rob , Mitchell; Heaton, Emily; Khanna, Madhu; Lee, DoKyoung (2025)

The compiled datasets include detailed costs for switchgrass production, categorized into establishment, maintenance, and harvesting expenses, along with revenue calculations. Costs were gathered from multiple sources and adjusted for inflation, focusing on farm-gate profitability, excluding fixed costs and transportation. All financial data is provided per hectare. The dataset was used to evaluate the economic performance of forage- and bioenergy-type switchgrass cultivars and their response to nitrogen fertilization across diverse marginal environments in the U.S. Midwest. Data Envelopment Analysis (DEA) and cost-benefit analysis were employed to assess the efficiency and profitability of 23 different cultivar and fertilization rate combinations over five years.

published: 2026-01-08

Convolutional Neural Network-based Sequence-to-Expression Prediction Tool (CoNSEPT)

Dibaeinia, Payam; Sinha, Saurabh (2026)

CoNSEPT is a tool to predict gene expression in various cis and trans contexts. Inputs to CoNSEPT are enhancer sequence, transcription factor levels in one or many trans conditions, TF motifs (PWMs), and any prior knowledge of TF-TF interactions.

keywords: software; gene expression

published: 2022-06-20

A Prototype Gutenberg-HathiTrust Sentence-level Parallel Corpus

Jiang, Ming; Dubnicek, Ryan; Worthey, Glen; Underwood, Ted; Downie, J. Stephen (2022)

This is a sentence-level parallel corpus in support of research on OCR quality. The source data comes from: (1) Project Gutenberg for human-proofread "clean" sentences; and, (2) HathiTrust Digital Library for the paired sentences with OCR errors. In total, this corpus contains 167,079 sentence pairs from 189 sampled books in four domains (i.e., agriculture, fiction, social science, world war history) published from 1793 to 1984. There are 36,337 sentences that have two OCR views paired with each clean version. In addition to sentence texts, this corpus also provides the location (i.e., sentence and chapter index) of each sentence in its belonging Gutenberg volume.

keywords: sentence-level parallel corpus; optical character recognition; OCR errors; Project Gutenberg; HathiTrust Digital Library; digital libraries; digital humanities;

published: 2022-04-29

Biological and Simulated datasets for testing the SCAMPP framework for phylogenetic placement methods

Wedell, Eleanor; Warnow, Tandy (2022)

Thank you for using these datasets! These files contain trees and reference alignments, as well as the selected query sequences for testing phylogenetic placement methods against and within the SCAMPP framework. There are four datasets from three different sources, each containing their source alignment and "true" tree, any estimated trees that may have been generated, and any re-estimated branch lengths that were created to be used with their requisite phylogenetic placement method. Three biological datasets (16S.B.ALL, PEWO/LTP_s128_SSU, and PEWO/green85) and one simulated dataset (nt78) is contained. See README.txt in each file for more information.

keywords: Phylogenetic Placement; Phylogenetics; Maximum Likelihood; pplacer; EPA-ng

published: 2025-08-27

Data for Identifying the best high-biomass sorghum hybrids based on biomass yield potential and feedstock quality affected by nitrogen fertility management under various environments

Jang, Chunhwa; Namoi, Nictor; Lee, Jung Woo; Becker, Talon; Rooney, William; Lee, DoKyoung (2025)

Data were collected from agronomy fields in Urbana and Ewing, IL, during the 2022 and 2023 growing seasons. The dataset includes dry biomass yield, nitrogen, phosphorus, and potassium concentrations and removals, and chemical composition elements (cellulose, hemicellulose, lignin, and soluble fractions) for 13 high-biomass sorghum hybrids. data_sharing.xlsx contains 20 columns and 104 rows. Below is the explanation of all variables in the file: Year: 2022; 2023 Location: Urbana, IL; Ewing, IL N rate (kg-N/ha): 0; 112 Hybrid #: H1-H13 Pedigree: Pedigree for 13 hybrids Dry biomass yield (Mg/ha): Aboveground dry biomass yield N (g/kg): Nitrogen concentration in plant tissue P (g/kg): Phosphorus concentration in plant tissue K (g/kg): Potassium concentration in plant tissue N (kg/ha): Nitrogen removal by aboveground biomass P (kg/ha): Phosphorus removal by aboveground biomass K (kg/ha): Potassium removal by aboveground biomass Cellulose (g/kg): Cellulose concentration in plant tissue Hemicellulose (g/kg): Hemicellulose concentration in plant tissue Lignin (g/kg): Lignin concentration in plant tissue Soluble (g/kg): Soluble concentration in plant tissue Cellulose (Mg/ha): Cellulose content in aboveground biomass Hemicellulose (Mg/ha): Hemicellulose content in aboveground biomass Lignin (Mg/ha): Lignin content in aboveground biomass Soluble (Mg/ha): Soluble content in aboveground biomass

keywords: high-biomass sorghum hybrids; yield potential; environmental adaptability; feedstock quality; nutrient removal; N fertilization

published: 2025-09-26

Data from Biodiesel Production from Engineered Sugarcane Lipids under Uncertain Feedstock Compositions: Process Design and Techno-Economic Analysis

Arora, Amit; Singh, Vijay (2025)

In this study, different process schemes were designed and evaluated for biodiesel production from engineered cane lipids with uncertain fatty acid compositions. Four different process schemes were compared under (i) thermal glycerolysis and (ii) enzymatic glycerolysis approaches. These schemes were based on the biodiesel yield and economic indicators such as the net present value (NPV) and the minimum selling price (MSP) of biodiesel. A scheme with polar lipid separation under thermal glycerolysis resulted in the maximum NPV ($96.5 million) and minimum MSP ($1107/ton biodiesel), respectively. Through local sensitivity analysis, it was concluded that the cane lipid percentage is the most significant factor influencing process economics. A conjoint analysis of the lipid procurement price and cane lipid percent suggested that 15% cane lipids with a low lipid procurement price ($0.536/kg) results in a positive NPV. When the cane lipid price is higher (>$0.80/kg), a 20% lipid content should be considered to achieve a positive NPV. At 20% cane lipids, the worst-case and best-case scenarios were evaluated by analyzing the interplay of the three most important parameters, The best-case scenario revealed that the minimum NPV under any process scheme could yield more than $100 million (or MSP: $0.80/L), and the worst-case analysis showed that losses incurred by the plant could be as high as $80 million (MSP: $1.36/L). A Monte Carlo simulation indicated that there is a 70% chance of the plant being profitable (NPV > 0).

keywords: Conversion;Economics;Feedstock Bioprocessing;Modeling

published: 2025-10-21

Data for Transformation and Gene Editing in the Bioenergy Grass Miscanthus

Trieu, Anthony; Belaffif, Mohammad B.; Hirannaiah, Pradeepa; Manjunatha, Shilpa; Wood, Rebekah; Bathula, Yokshitha; Billingsley, Rebecca L.; Arpan, Anjali; Sacks, Erik; Clemente, Tom; Moose, Stephen; Reichert, Nancy A.; swaminathan, kankshita (2025)

Miscanthus, a C4 member of the family Poaceae, is a promising perennial crop for bioenergy, renewable bioproducts, and carbon sequestration. Species of interest include nothospecies Miscanthus x giganteus and its parental species M. sacchariflorus and M. sinensis. Use of biotechnology-based procedures to genetically improve miscanthus, to date, have only included plant transformation procedures for introduction of exogenous genes into the host genome at random, non-targeted sites.

keywords: Feedstock Production;Biomass Analytics;Genomics

published: 2023-07-14

Pollen of Podocarpus (Podocarpaceae): Airyscan confocal superresolution images

Punyasena, Surangi W.; Urban, Michael A.; Adaime, Marc-Elie; Romero, Ingrid; Jaramillo, Carlos (2023)

This dataset includes a total of 300 images of 45 extant species of Podocarpus (Podocarpaceae) and nine images of fossil specimens of the morphogenus Podocarpidites. The goal of this dataset is to capture the diversity of morphology within the genus and create an image database for training machine learning models. The images were taken using Airyscan confocal superresolution microscopy at 630x magnification (63x/NA 1.4 oil DIC). The images are in the CZI file format. They can be opened using Zeiss propriety software (Zen, Zen lite) or open microscopy software, such as ImageJ. More information on how to open CZI files can be found here: [https://www.zeiss.com/microscopy/us/products/software/zeiss-zen/czi-image-file-format.html] Please cite this dataset and listed publications when using these images.

keywords: optical superresolution microscopy; Zeiss Airyscan; CZI images; conifer; saccate pollen; Podocarpus; Podocarpidites; Smithsonian Tropical Research Institute

published: 2019-05-31

Frequent pattern subject transactions from the University of Illinois Library (2016 - 2018)

Hahn, Jim (2019)

The data are provided to illustrate methods in evaluating systematic transactional data reuse in machine learning. A library account-based recommender system was developed using machine learning processing over transactional data of 383,828 transactions (or check-outs) sourced from a large multi-unit research library. The machine learning process utilized the FP-growth algorithm over the subject metadata associated with physical items that were checked-out together in the library. The purpose of this research is to evaluate the results of systematic transactional data reuse in machine learning. The analysis herein contains a large-scale network visualization of 180,441 subject association rules and corresponding node metrics.

keywords: evaluating machine learning; network science; FP-growth; WEKA; Gephi; personalization; recommender systems

published: 2020-06-26

Data from: Quantifying Errors in the Aerosol Mixing-State Index Based on Limited Particle Sample Size

Gasparik, Jessica T.; Ye, Qing; Curtis, Jeffrey H.; Presto, Albert A.; Donahue, Neil M.; Sullivan, Ryan C.; West, Matthew; Riemer, Nicole (2020)

This dataset contains the PartMC-MOSAIC simulations used in the article "Quantifying Errors in the Aerosol Mixing-State Index Based on Limited Particle Sample Size". The 1000 simulations of output data is organized into a series of archived folders, each containing 100 scenarios. Within each scenario directory are 25 NetCDF files, which are the hourly output of a PartMC-MOSAIC simulation containing all information regarding the environment, particle and gas state. This dataset was used to investigate the impact of sample size on determining aerosol mixing state. This data may be useful as a data set for applying different types of estimators.

keywords: Atmospheric aerosols; single-particle measurements; sampling uncertainty; NetCDF

published: 2022-07-19

Effect of Micro-patterned Mucin on Quinolone and Rhamnolipid Profiles of Mucoid Pseudomonas aeruginosa under Antibiotic Stress

Parmar, Dharmeshkumar; Jia, Jin; Shrout, Joshua; Sweedler, Jonathan; Bohn, Paul (2022)

#### Details of Pseudomonas aeruginosa biofilm dataset #### ----------------*Folder Structure*------------------------------------- This dataset contains peak intensity tables extracted from mass spectrometry imaging (MSI) data using tools, SCiLS and MSI reader. There are 2 folders in "MSI-Data-Paeruginosa-biofilms-UIUC-DP-JVS-July2022.zip", each folder contains 3 sub-folders as listed below. 1. PellicleBiofilms-and-Supernatant [Pellicle biofilms collected from air-liquid interface and spend supernatant medium after 96 h incubation period]: (1) Full-Scan-Data-96h; (2) MSMS-data-from-C7-Quinolones-96h; and (3) MSMS-data-from-C9-Quinolones-96h 2. StaticBiofilms [Static biofilms grown on mucin surface]: (1) Full-Scan-Data; (2) MSMS-data-from-C7-Quinolones; and (3) MSMS-data-from-C9-Quinolones ----------------*File name*---------------------------------------------- Sample information is included in the file names for easy identification and processing. Attributes covered in file names are explained in the example below. *Example file name "Rep1-Stat-FRD1-mPat-48-FS"* ~ Each unit of information is separated by "-" ~Unit 1 - "Rep1" - Biological replicate ( Rep1, Rep2, and Rep3) ~Unit 2 - "Stat" - Sample type (Stat = Static Biofilm, Pel = Pellicle biofilm, Sup = Supernatant) ~Unit 3 - "FRD1" - Strain (FRD1 = Mucoid strain, PAO1C = Non-mucoid strain) ~Unit 4 - "mPat" - Type of mucin surface used (mPat = patterned mucin surface, mUni = uniform mucin surface) ~Unit 5 - "48" - Sample time point (hours = 48, 72, 96) ~Unit 6 - "FS" - Scan type used in MSI (FS = high resolution full-scan, 260 = targeted MS/MS of C7 quinolones (m/z 260), 288 = targeted MS/MS of C9 quinolones (m/z 288)) ----------------*File structure*------------------------------------------ All MSI data has been exported to CSV format. Each CSV files contains information about scan number, Coordinates (x,y,z), m/z values, extraction window (absolute), and corresponding intensities in the form of a matrix. ----------------*End of Information*--------------------------------------

keywords: mass spectrometry imaging (MSI); biofilm; antibiotic resistance; Pseudomonas aeruginosa; quorum sensing; rhamnolipids

published: 2025-10-10

Data for Robust Paths to Net Greenhouse Gas Mitigation and Negative Emissions via Advanced Biofuels

Field, John L.; Richard, Tom; Smithwick, Erica A. H.; Cai, Hao; Laser, Mark; LeBauer, David; Long, Stephen; Paustian, Keith; Qin, Zhangcai; Sheehan, John; Smith, Pete; Wang, Michael Q.; Lynd, Lee (2025)

This zip file contains a UNIX-format DayCent model executable, input files, automation code, and associated directory structure necessary to re-produce the DayCent analysis underlying the manuscript. The main script “autodaycent.py” (written for Python 2.7) opens an interactive command line routine that facilitates: Calibrating the DayCent pine growth model; Initializing DayCent for a set of case studies sites; Executing an ensemble of model runs representing case study site reforestation, grassland restoration, or conversion to switchgrass cultivation; and Results analysis & generation of manuscript Fig. 3. Note that the interactive analysis code requires that all input files to be contained in the directory structure as uploaded, without modification. Executable versions of the DayCent model compatible with other operating systems are available upon request.

keywords: Feedstock Production;Modeling

published: 2025-11-24

Data for A Transcriptomic Atlas of Acute Stress Response to Low pH in Multiple Issatchenkia orientalis Strains

Dubinkina, Veronika; Bhogale, Shounak; Hsieh, Ping-Hung; Dibaeinia, Payam; Nambiar, Ananthan; Maslov, Sergei; Yoshikuni, Yasuo; Sinha, Saurabh (2025)

Because of its natural stress tolerance to low pH, Issatchenkia orientalis (a.k.a. Pichia kudriavzevii) is a promising non-model yeast for bio-based production of organic acids. Yet, this organism is relatively unstudied, and specific mechanisms of its tolerance to low pH are poorly understood, limiting commercial use. In this study, we selected 12 I. orientalis strains with varying acid stress tolerance (six tolerant and six susceptible) and profiled their transcriptomes in different pH conditions to study potential mechanisms of pH tolerance in this species. We identified hundreds of genes whose expression response is shared by tolerant strains but not by susceptible strains, or vice versa, as well as genes whose responses are reversed between tolerant and susceptible strains. We mapped regulatory mechanisms of transcriptomic responses via motif analysis as well as differential network reconstruction, identifying several transcription factors, including Stb5, Mac1, and Rtg1/Rtg3, some of which are known for their roles in acid response in Saccharomyces cerevisiae. Functional genomics analysis of short-listed genes and transcription factors suggested significant roles for energy metabolism and translation-related processes, as well as the cell wall integrity pathway and RTG-dependent retrograde signaling pathway. Finally, we conducted additional experiments for two organic acids, 3-hydroxypropionate and citramalate, to eliminate acid-specific effects and found potential roles for glycolysis and trehalose biosynthesis specifically for response to low pH. In summary, our approach of comparative transcriptomics and phenotypic contrasting, along with a multi-pronged bioinformatics analysis, suggests specific mechanisms of tolerance to low pH in I. orientalis that merit further validation through experimental perturbation and engineering.

keywords: Conversion;Transcriptomics

published: 2022-11-11

Data for Chemical Short-Range Ordering in a CrCoNi Medium-Entropy Alloy

Hsiao, Haw-Wen; Zuo, Jian-Min (2022)

This dataset is for characterizing chemical short-range-ordering in CrCoNi medium entropy alloys. It has three sub-folders: 1. code, 2. sample WQ, 3. sample HT. The software needed to run the files is Gatan Microscopy Suite® (GMS). Please follow the instruction on this page to install the DM3 GMS: <a href="https://www.gatan.com/installation-instructions#Step1">https://www.gatan.com/installation-instructions#Step1</a> 1. Code folder contains three DM scripts to be installed in Gatan DigitalMicrograph software to analyze scanning electron nanobeam diffraction (SEND) dataset: Cepstrum.s: need [EF-SEND_sampleWQ_cropped_aligned.dm3] in Sample WQ and the average image from [EF-SEND_sampleWQ_cropped_aligned.dm3]. Same for Sample HT folder. log_BraggRemoval.s: same as above. Patterson.s: Need refined diffuse patterns in Sample HT folder. 2. Sample WQ and 3. Sample HT folders both contain the SEND data (.ser) and the binned SEND data (.dm3) as well as our calculated strain maps as the strain measurement reference. The Sample WQ folder additionally has atomic resolution STEM images; the Sample HT folder additionally has three refined diffuse patterns as references for diffraction data processing. * Only .ser file is needed to perform the strain measurement using imToolBox as listed in the manuscript. .emi file contains the meta data of the microscope, which can be opened together with .ser file using FEI TIA software.

keywords: Medium entropy alloy; CrCoNi; chemical short-range-ordering; CSRO; TEM

published: 2024-02-21

Data for "Niche conservatism and spread explain hybridization and introgression between native and invasive fish"

Hartman, Jordan H; Corush, Joel B; Larson, Eric R; Tiemann, Jeremy S; Willink, Philip; Davis, Mark A (2024)

Data associated with the manuscript "Niche conservatism and spread explain hybridization and introgression between native and invasive fish" by Jordan H. Hartman, Joel B. Corush, Eric R. Larson, Jeremy S. Tiemann, Philip Willink, and Mark A. Davis. For this project, we combined results of ecological niche models (ENMs) and next-generation restriction site-associated DNA sequencing (RADseq) to test theories of niche conservatism and biotic resistance on the success of invasion, hybridization, and extent of introgression between native Western Banded Killifish and non-native Eastern Banded Killifish. This dataset provides the sampling locations and number of Banded Killifish in each population, accession numbers for RADseq from the National Center for Biotechnology Information Sequence Read Archive and the assignment of each Banded Killifish, the habitat associations of each population from the ENMs, and the occurrence points used to build the ENMs.

keywords: Banded Killifish; ecological niche model; Fundulus diaphanus; hybrid swarm; invasive species; Laurentian Great Lakes

published: 2024-03-28

Enhancing Carrier Mobility In Monolayer MoS2 Transistors With Process induced Strain

Zhang, Yue; Zhao, Helin; Huang, Siyuan; Hossain, Mohhamad Abir; van der Zande, Arend (2024)

Read me file for the data repository ******************************************************************************* This repository has raw data for the publication "Enhancing Carrier Mobility In Monolayer MoS2 Transistors With Process Induced Strain". We arrange the data following the figure in which it first appeared. For all electrical transfer measurement, we provide the up-sweep and down-sweep data, with voltage units in V and conductance unit in S. All Raman modes have unit of cm^-1. ******************************************************************************* How to use this dataset All data in this dataset is stored in binary Numpy array format as .npy file. To read a .npy file: use the Numpy module of the python language, and use np.load() command. Example: suppose the filename is example_data.npy. To load it into a python program, open a Jupyter notebook, or in the python program, run: import numpy as np data = np.load("example_data.npy") Then the example file is stored in the data object. *******************************************************************************

published: 2025-11-06

Data for Photoenzymatic Asymmetric Hydroamination for Chiral Alkyl Amine Synthesis

Harrison, Wesley; Jiang, Guangde; Zhang, Zhengyi; Li, Maolin; Chen, Haoyu; Zhao, Huimin (2025)

Chiral alkyl amines are common structural motifs in pharmaceuticals, natural products, synthetic intermediates, and bioactive molecules. An attractive method to prepare these molecules is the asymmetric radical hydroamination; however, this approach has not been explored with dialkyl amine-derived nitrogen-centered radicals since designing a catalytic system to generate the aminium radical cation, to suppress deleterious side reactions such as α-deprotonation and H atom abstraction, and to facilitate enantioselective hydrogen atom transfer is a formidable task. Herein, we describe the application of photoenzymatic catalysis to generate and harness the aminium radical cation for asymmetric intermolecular hydroamination. In this reaction, the flavin-dependent ene-reductase photocatalytically generates the aminium radical cation from the corresponding hydroxylamine and catalyzes the asymmetric intermolecular hydroamination to furnish the enantioenriched tertiary amine, whereby enantioinduction occurs through enzyme-mediated hydrogen atom transfer. This work highlights the use of photoenzymatic catalysis to generate and control highly reactive radical intermediates for asymmetric synthesis, addressing a long-standing challenge in chemical synthesis.

keywords: Conversion;Bioproducts;Catalysis

published: 2025-12-01

Data for "Modeling the Global Citation Network using the Scalable Agent-based Simulator for Citation Analysis with Recency-emphasized Sampling (SASCA-ReS)"

Park, Minhyuk; Yi, Haotian; Warnow, Tandy; Chacko, George (2025)

This dataset principally consists of four synthetic citation networks that were generated during the preparation of the manuscript Park M, Yi H, Warnow T, and Chacko G (2025). Modeling the Global Citation Network using the Scalable Agent-based Simulator for Citation Analysis with Recency-emphasized Sampling (SASCA-ReS). A preprint is available on Zenodo (below) and the manuscript has been submitted to the MetaRoR platform for review and feedback. @misc{park_2025_17789558, author = {Park, Minhyuk and Yi, Haotian and Warnow, Tandy and Chacko, George}, title = {Modeling the Global Citation Network using the Scalable Agent-based Simulator for Citation Analysis with Recency-emphasized Sampling (SASCA- ReS) }, month = dec, year = 2025, publisher = {Zenodo}, doi = {10.5281/zenodo.17789558}, url = {https://doi.org/10.5281/zenodo.17789558}, } The networks are roughly 14, 76, 161, and 218 million nodes each. Both nodelists with attributes and edge lists are provided as gzipped parquet files along with the configuration file that was passed to the SASCA-ReS software, which can be accessed at: <a href="https://github.com/illinois-or-research-analytics/SASCA-ReS">https://github.com/illinois-or-research-analytics/SASCA-ReS</a>. A copy of the configuration file that was used to generate the network with SASCA-ReS is also provided. For example: abm14_config.ini; abm14_edgelist.parquet.gz; and abm14_nodelist.parquet.gz. The column headers in the edgelists and nodelists and the fields in the configuration file are explained in the Github repository for SASCA-ReS. In addition, we provide sj_reccount, a table of real world citation frequencies that is an input to the SASCA-Res software. The first column (diff) of sj_reccount lists the difference between the publication year of a citing document and the publication year of a cited document. The second column (count) reports the frequency of such citations across the dataset of 77879427 observations, which is derived from the biomedical literature. Finally, we share data, composite_maverick_disruption.csv , from the mavericks (unconventional citing strategies) experiment reported in the Park et al. (2025) manuscript available at <a href="https://zenodo.org/records/17772113">https://zenodo.org/records/17772113</a>. The columns in the composite_maverick_disruption.csv file are: node_id -> of agents in the various simulations n_i, n_j, n_k -> terms used to compute disruption per "Wu, L., Wang, D. & Evans, J.A. Large teams develop and small teams disrupt science and technology. Nature 566, 378–382 (2019). <a href="https://doi.org/10.1038/s41586-019-0941-9">https://doi.org/10.1038/s41586-019-0941-9"</a> disruption -> the disruption metric of Wu, Wang, and Evans (2019) type -> maverick type (maximizer, randomnik, or minimizer) year -> virtual year in the simulation when the maverick was created alpha -> the alpha parameter of the control agent pa_weight -> the preferential attachment weight of the control agent phenotype fit_peak_value -> the fitness value assigned to the control agent in_degree -> the count of citations accumulated by the maverick or control agent at the end of the simulation out_degree -> the count of references made by the maverick tag -> a label for the experiment, e.g. od249_f1 indicates that the mavericks in this experiment made 249 citations and were assigned a fitness value of 1.

keywords: synthetic networks; agent based models; SASCA-ReS; citation networks