Science Research Source Guide: Big Data, Machine Learning, and Computational Systems
A comprehensive guide to understanding big data systems, machine learning pitfalls in healthcare, and computational optimization frameworks using sources [1]–[4].
Question-ready source guide
Djoomba source guide · Start with the evidence
Automatically generated by Djoomba using Qwen3-8B. Not peer reviewed. Read and cite the underlying studies below.
Key findings
- Big data systems require specialized architectures for real-time analysis and storage [1].
- Machine learning models for COVID-19 detection face methodological flaws and biases [2].
- Radio-computational resource optimization reduces energy consumption in mobile-edge computing [3].
- DScribe provides standardized descriptors for machine learning in materials science [4].
Frame the question
This guide addresses how emerging technologies like big data systems, machine learning, and computational optimization address modern scientific challenges. It explores the interplay between data processing frameworks (e.g., Hadoop [1]), healthcare diagnostics (e.g., ML for COVID-19 [2]), and materials science (e.g., DScribe [4]). The sources collectively highlight the need for robust methodologies, ethical considerations, and interdisciplinary approaches to solve complex problems.
What the evidence shows
The anchor source [1] defines big data as 'data from distinctive domains... requiring real-time analysis' and introduces the Hadoop framework as a solution. Source [2] critiques ML models for COVID-19 detection, noting that 'none of the models identified are of potential clinical use due to methodological flaws.' Source [3] demonstrates how joint radio-computational optimization 'minimizes overall users' energy consumption.' Source [4] emphasizes DScribe's role in 'accelerating machine learning for atomistic property prediction' through descriptors like Coulomb matrices and SOAP. These sources collectively show how data challenges, from healthcare to materials science, demand tailored solutions.
Follow the source trail
Source [1] establishes the foundational framework for big data systems, which [4] extends to materials science by providing descriptors for ML. Source [2] contrasts with [1] by highlighting the limitations of ML in healthcare, while [3] bridges computational and radio resource management. The Hadoop framework in [1] is distinct from DScribe's descriptors in [4], but both address data processing needs. Source [2] and [3] share a focus on optimization but differ in application: [2] critiques ML models, while [3] proposes a solution for energy efficiency. All sources converge on the theme of data-driven innovation with methodological rigor.
Use these sources well
For an essay on big data systems, use [1] to explain the 'four sequential modules' of data value chains and Hadoop's role. Cite [4] to contrast big data frameworks with materials science ML tools. When discussing healthcare ML, reference [2] to highlight methodological flaws and its recommendations for 'higher-quality model development.' For computational optimization, use [3] to detail the 'successive convex approximation technique' and its 'distributed implementation.' Compare [1] and [4] to show how data processing frameworks evolve across domains. Always pair [2] with [1] to contextualize ML limitations within broader data challenges.
What to search next
How might the Hadoop framework in [1] address the scalability issues faced by DScribe in [4]? Can the methodological critiques in [2] inform improvements in big data analytics platforms described in [1]? What ethical considerations arise when applying DScribe's descriptors to sensitive materials science data? How could the 'successive convex approximation' technique in [3] be adapted to optimize ML models for healthcare diagnostics? These questions suggest pathways for deeper analysis, such as exploring cross-domain optimization strategies or evaluating the reproducibility of ML models in [2] against big data benchmarks in [1].
Verbatim source abstracts
[1] Toward Scalable Systems for Big Data Analytics: A Technology Tutorial — IEEE Access, 2014-01-01, doi:10.1109/access.2014.2332453
Recent technological advancements have led to a deluge of data from distinctive domains (e.g., health care and scientific sensors, user-generated data, Internet and financial companies, and supply chain systems) over the past two decades. The term big data was coined to capture the meaning of this emerging trend. In addition to its sheer volume, big data also exhibits other unique characteristics as compared with traditional data. For instance, big data is commonly unstructured and require more real-time analysis. This development calls for new system architectures for data acquisition, transmission, storage, and large-scale data processing mechanisms. In this paper, we present a literature survey and system tutorial for big data analytics platforms, aiming to provide an overall picture for nonexpert readers and instill a do-it-yourself spirit for advanced audiences to customize their own big-data solutions. First, we present the definition of big data and discuss big data challenges. Next, we present a systematic framework to decompose big data systems into four sequential modules, namely data generation, data acquisition, data storage, and data analytics. These four modules form a big data value chain. Following that, we present a detailed survey of numerous approaches and mechanisms from research and industry communities. In addition, we present the prevalent Hadoop framework for addressing big data challenges. Finally, we outline several evaluation benchmarks and potential research directions for big data systems. [1]
[2] Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans — Nature Machine Intelligence, 2021-03-15, doi:10.1038/s42256-021-00307-0
Abstract Machine learning methods offer great promise for fast and accurate detection and prognostication of coronavirus disease 2019 (COVID-19) from standard-of-care chest radiographs (CXR) and chest computed tomography (CT) images. Many articles have been published in 2020 describing new machine learning-based models for both of these tasks, but it is unclear which are of potential clinical utility. In this systematic review, we consider all published papers and preprints, for the period from 1 January 2020 to 3 October 2020, which describe new machine learning models for the diagnosis or prognosis of COVID-19 from CXR or CT images. All manuscripts uploaded to bioRxiv, medRxiv and arXiv along with all entries in EMBASE and MEDLINE in this timeframe are considered. Our search identified 2,212 studies, of which 415 were included after initial screening and, after quality screening, 62 studies were included in this systematic review. Our review finds that none of the models identified are of potential clinical use due to methodological flaws and/or underlying biases. This is a major weakness, given the urgency with which validated COVID-19 models are needed. To address this, we give many recommendations which, if followed, will solve these issues and lead to higher-quality model development and well-documented manuscripts. [2]
[3] Joint Optimization of Radio and Computational Resources for Multicell Mobile-Edge Computing — IEEE Transactions on Signal and Information Processing over Networks, 2015-06-01, doi:10.1109/tsipn.2015.2448520
Migrating computational intensive tasks from mobile devices to more resourceful cloud servers is a promising technique to increase the computational capacity of mobile devices while saving their battery energy. In this paper, we consider an MIMO multicell system where multiple mobile users (MUs) ask for computation offloading to a common cloud server. We formulate the offloading problem as the joint optimization of the radio resources-the transmit precoding matrices of the MUs-and the computational resources-the CPU cycles/second assigned by the cloud to each MU-in order to minimize the overall users' energy consumption, while meeting latency constraints. The resulting optimization problem is nonconvex (in the objective function and constraints). Nevertheless, in the single-user case, we are able to compute the global optimal solution in closed form. In the more challenging multiuser scenario, we propose an iterative algorithm, based on a novel successive convex approximation technique, converging to a local optimal solution of the original nonconvex problem. We then show that the proposed algorithmic framework naturally leads to a distributed and parallel implementation across the radio access points, requiring only a limited coordination/signaling with the cloud. Numerical results show that the proposed schemes outperform disjoint optimization algorithms. [3]
[4] DScribe: Library of descriptors for machine learning in materials science — Computer Physics Communications, 2019-09-26, doi:10.1016/j.cpc.2019.106949
DScribe is a software package for machine learning that provides popular feature transformations (“descriptors”) for atomistic materials simulations. DScribe accelerates the application of machine learning for atomistic property prediction by providing user-friendly, off-the-shelf descriptor implementations. The package currently contains implementations for Coulomb matrix, Ewald sum matrix, sine matrix, Many-body Tensor Representation (MBTR), Atom-centered Symmetry Function (ACSF) and Smooth Overlap of Atomic Positions (SOAP). Usage of the package is illustrated for two different applications: formation energy prediction for solids and ionic charge prediction for atoms in organic molecules. The package is freely available under the open-source Apache License 2.0. Program Title: DScribe Program Files doi: http://dx.doi.org/10.17632/vzrs8n8pk6.1 Licensing provisions: Apache-2.0 Programming language: Python/C/C++ Supplementary material: Supplementary Information as PDF Nature of problem: The application of machine learning for materials science is hindered by the lack of consistent software implementations for feature transformations. These feature transformations, also called descriptors, are a key step in building machine learning models for property prediction in materials science. Solution method: We have developed a library for creating common descriptors used in machine learning applied to materials science. We provide an implementation the following descriptors: Coulomb matrix, Ewald sum matrix, sine matrix, Many-body Tensor Representation (MBTR), Atom-centered Symmetry Functions (ACSF) and Smooth Overlap of Atomic Positions (SOAP). The library has a python interface with computationally intensive routines written in C or C++. The source code, tutorials and documentation are provided online. A continuous integration mechanism is set up to automatically run a series of regression tests and check code coverage when the codebase is updated. [4]
Limitations
- Source [2] excludes post-October 2020 studies, limiting its scope.
- Source [4] focuses on materials science, which may not address broader ML challenges.
- Source [3] assumes single-user scenarios for simplicity, which may not reflect real-world multicell systems.
- Source [1] emphasizes system architecture but does not address ethical implications of big data.
Underlying research
Sources and citation tools
Copy a citation for the original publication—not a fabricated Djoomba author. Numbering matches the markers in this source guide.
Source 1 · Anchor
Toward Scalable Systems for Big Data Analytics: A Technology Tutorial
Han Hu, Yonggang Wen, Tat‐Seng Chua, Xuelong Li · IEEE Access · 2014
Source 2
Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans
Michael Roberts, Derek Driggs, Matthew Thorpe, Julian Gilbey, Michael Yeung, Stephan Ursprung, Angelica I. Aviles-Rivero, Christian Etmann, Cathal McCague, Lucian Beer, Jonathan R. Weir-McCall, Zhongzhao Teng, Effrossyni Gkrania-Klotsas, AIX-COVNET, Alessandro Ruggiero, Anna Korhonen, Emily Jefferson, Emmanuel Ako, Georg Langs, Ghassem Gozaliasl, Guang Yang, Helmut Prosch, Jacobus Preller, Jan Stanczuk, Jing Tang, Johannes Hofmanninger, Judith Babar, Lorena Escudero Sánchez, Muhunthan Thillai, Paula Martin Gonzalez, Philip Teare, Xiaoxiang Zhu, Mishal Patel, Conor Cafolla, Hojjat Azadbakht, Joseph Jacob, Josh Lowe, Kang Zhang, Kyle Bradley, Marcel Wassin, Markus Holzer, Kangyu Ji, Maria Delgado Ortet, Tao Ai, Nicholas Walton, Pietro Lio, Samuel Stranks, Tolou Shadbahr, Weizhe Lin, Yunfei Zha, Zhangming Niu, James H. F. Rudd, Evis Sala, Carola-Bibiane Schönlieb · Nature Machine Intelligence · 2021
Source 3
Joint Optimization of Radio and Computational Resources for Multicell Mobile-Edge Computing
Stefania Sardellitti, Gesualdo Scutari, Sergio Barbarossa · IEEE Transactions on Signal and Information Processing over Networks · 2015
Source 4
DScribe: Library of descriptors for machine learning in materials science
Lauri Himanen, Marc O. J. Jäger, Eiaki V. Morooka, Filippo Federici Canova, Yashasvi S. Ranawat, David Gao, Patrick Rinke, Adam S. Foster · Computer Physics Communications · 2019