Home Browse Just accepted

Online First

Accepted, unedited articles published online and citable. The final edited and typeset version of record will appear in the future.
Please wait a minute...
  • Select all
    |
  • Research Article
    Chao Tang, Yan Qi, Haiyun Xu, Shuying Li, Zenghui Yue, Robin Haunschild
    Journal of Data and Information Science. https://doi.org/10.1515/jdis-2025-0120
    Accepted: 2026-07-30
    Abstract
    Purpose

    As technological competition increasingly becomes the focus of global attention, identifying key core technologies (KCTs) is of great significance for seizing the technological high ground, promoting national economic development, and safeguarding national security.

    Design/methodology/approach

    Based on a clear definition of the concept and characteristics of KCTs, this study proposes a field–topic–technology (FTT) identification method, which follows a progressive, multi-method fusion approach that combines primary and auxiliary recognition paths. The FTT method is empirically applied to patents in the field of small-molecule targeted therapies for lung cancer. The primary path of the method proceeds as “core subfield → core topic → KCTs”, while the auxiliary path follows “core subfield → stage-based topic evolution trends → auxiliary identification of KCTs.” Through the combination of primary and auxiliary paths and a progressive hierarchical approach, this method achieves systematic identification of KCTs. Meanwhile, the LightGBM model is employed for indicator weighting, enhancing the scientific validity and objectivity of the weighting process.

    Findings

    The study successfully identifies six KCTs in the field of small-molecule targeted therapies for lung cancer. Comparative analyses with other identification methods as well as with official documents and high-impact publications verify the scientificity, feasibility, effectiveness, and robustness of the FTT method. Empirical evidence demonstrates that this method is suitable for fields with abundant patent data and relatively clear technological evolution pathways.

    Research limitations

    Although this study constructs a KCTs identification framework, several limitations remain in the identification process. First, the characteristics of KCTs are multidimensional and complex; this study selects indicators only from four dimensions, which makes it difficult to fully capture their intrinsic attributes. Second, the constructed indicator system mainly relies on quantitative indicators, with insufficient representation of qualitative factors such as tacit knowledge and engineering experience. Third, this study is primarily based on patent data and does not fully integrate multi-source information such as academic publications, industry standards, and policy documents, which may lead to certain biases in the identification results. In addition, the empirical validation is conducted in a single field, and the cross-field applicability of the method still requires further examination. Future research may enhance the comprehensiveness and robustness of KCTs identification by expanding feature dimensions, integrating multi-source heterogeneous data, and conducting cross-field validation.

    Practical implications

    The proposed FTT approach supports technology intelligence and strategic decision-making through KCT identification and technological evolution tracking. It can be extended to other patent‑rich technology fields to support R&D planning and policy formulation.

    Originality/value

    Addressing existing research gaps – such as limited conceptual dimensions, fragmented identification methods, and lack of focus – this study defines the concept and attribute features of KCTs from multiple dimensions and constructs a hierarchical, multi-dimensional, and multi-method KCT identification framework that progresses from the macro field level to the micro technology level, integrating both primary and auxiliary recognition paths. Compared with previous studies that relied on single methods or linear processes, the proposed approach demonstrates stronger logical hierarchy, systematic structure, and methodological rigor. Furthermore, by introducing a machine learning-based approach for dynamic learning and nonlinear optimization of indicator weights, this study overcomes the limitations of traditional weighting methods and enhances the scientificity and objectivity of the identification results. Overall, this research achieves conceptual and methodological innovation in KCT identification, offering new perspectives for hierarchical system construction and indicator weighting optimization, and providing valuable references for subsequent studies and practical applications.

  • Research Article
    Arielle J. King, Sayed A. Mostafa
    Journal of Data and Information Science. https://doi.org/10.1515/jdis-2026-0026
    Accepted: 2026-07-21
    Abstract
    Purpose

    This study examines how data-driven research (DDR) has diffused across U.S. higher education institutions and investigates its relationship with scientific impact as measured by citation counts. It aims to clarify how institutional context, disciplinary affiliation, and publication characteristics shape the visibility of data-intensive research.

    Design/methodology/approach

    Using publications indexed in the Web of Science from 2013 to 2023, DDR is identified through a transparent text-mining approach based on abstract-level keywords related to artificial intelligence, machine learning, big data, and data science. Disciplinary, institutional, and demographic patterns in DDR output are analyzed. Citation counts are modeled using Zero-Inflated Negative Binomial regression, Random Forest, eXtreme Gradient Boosting, and Support Vector Regression. Feature variables capture DDR status, institutional classification, population served, and publication-level characteristics.

    Findings

    Results indicate that DDR output expanded across all research areas, with the strongest growth in computer science and engineering and disproportionately large contributions from R1 institutions. Across modeling approaches, publication age, number of authors, and DDR status consistently emerge as the most influential predictors of citation impact.

    Research limitations

    DDR identification relies on keyword-based text classification, which may omit relevant studies or include unrelated work. The Web of Science database emphasizes English-language and high-impact journals, potentially limiting coverage. In addition, the observational design supports association rather than causal inference.

    Practical implications

    The findings inform research evaluators, policymakers, and institutional leaders about patterns of methodological diffusion and citation visibility. The proposed framework can support evidence-based assessment of data-driven scholarship and guide capacity-building efforts, particularly in less-resourced institutions.

    Originality/value

    This study integrates bibliometric analysis, text mining, and machine-learning methods to examine data-driven research at scale. It provides a reproducible approach for identifying DDR and offers new empirical evidence on how institutional and disciplinary contexts shape research visibility, contributing to science-of-science and research evaluation literature.

  • Research Article
    Jiajia Liu, Hanyue Sun, Robin Haunschild
    Journal of Data and Information Science. https://doi.org/10.1515/jdis-2025-0464
    Accepted: 2026-07-21
    Abstract
    Purpose

    Research on inverse problems of geophysics is extensive. Thus, there is a need to conduct in-depth and comprehensive bibliometric research to capture the overall development of the field.

    Design/methodology/approach

    We conduct a bibliometric analysis using CiteSpace to reveal dynamic trends in collaboration, co-citation, and co-occurrence in this field.

    Findings

    We find that publications in the field have shown a fluctuating upward trend since 2004, accompanied by a gradual increase in its proportion within the broader geophysics’ literature. In terms of institutional collaborations, Chinese and French research institutions occupy an important position in the global collaboration network. From a co-citation perspective, Geophysics is the most cited journal, and Araya-Polo Mauricio’s article (Araya-Polo, M., J. Jennings, A. Adler, and T. Dahlke. 2018. “Deep-Learning Tomography.” The Leading Edge 37 (1): 58–66) is the most cited reference. From the perspective of keyword co-occurrence, the keyword clusters show the diversity of geophysical inversion research. The research topics cover fundamental theory, method development, and practical applications. In addition, future research may focus more on research based on data-driven workflows and artificial intelligence.

    Research limitations

    Only the SCI and SSCI editions in the Web of Science database were used as the data source, ignoring the literature in other databases. We focused on the literature from 2004 to 2023, which excludes the latest literature in 2024 and earlier literature before 2004.

    Practical implications

    We provided a comprehensive and specific analysis of the literature in the field over the last 20 years to help readers gain a comprehensive understanding of the current state of research, hotspots, evolution, and trends in the field. We discussed current research challenges and possible future research directions to help researchers identify appropriate research directions and conduct subsequent research more efficiently.

    Originality/value

    Compared with previous related studies, this study is innovative as we provide a comprehensive and specific analysis of the literature in the field over the last 20 years to help readers gain a comprehensive understanding of the current state of research, hotspots, evolution, and trends in the field and discuss current research challenges and possible future research directions to help researchers identify appropriate research directions and conduct subsequent research more efficiently.

  • Research Article
    Jing Chen, Youjin Shi
    Journal of Data and Information Science. https://doi.org/10.1515/jdis-2025-0487
    Accepted: 2026-07-09
    Abstract
    Purpose

    This study explores the patterns of critical thinking in human-AI co-creation tasks and how they relate to creative performance, aiming to clarify the relationships between these patterns and creative performance and deepen understanding of the cognitive processes.

    Design/methodology/approach

    The dataset comes from human-AI dialogue data for two level creativity tasks, text summarization and creative writing, on the Hugging Face platform. A critical thinking coding scheme is developed to analyze critical thinking skills, then lag sequential analysis (LSA) and frequent sequence mining (FSM) are used to identify the critical thinking patterns. Statistical methods are employed to explore associations between these patterns and creative performance.

    Findings

    Six task-dependent critical thinking patterns are identified in human-AI co-creation: Explanatory consolidation (EC), analytical structuring (AS), evaluative specification (ES), evaluative decomposition (ED), evaluative regulation (ER) and inferential self-loop (IS). Whereas low and medium-order thinking skills (patterns EC, AS, ES) dominated the low-creativity task, the high-creativity task incorporated low, medium and high-order skills (patterns ES, ED, ER, and IS). Furthermore, these patterns show associations with creative performance in the creative writing task: the ER pattern supports comprehensive creativity, the ED and ES patterns facilitate partial creativity, the IS pattern acts as a limiting factor.

    Research limitations

    Reliance on merged large-scale datasets analyses may constrain interpretation of human-AI co-creation, and future studies should validate these findings through experimental approaches and single-dataset replication.

    Practical implications

    The findings can help educators design tasks at different levels of creativity to elicit students’ critical thinking, and help writers design personalized prompts that enhance specific dimensions of creativity (originality or elaboration) in line with their goals.

    Originality/value

    By adapting frameworks of critical thinking to human-AI co-creation, the study contributes to a comprehensive understanding of critical thinking patterns, compares these patterns across different creative tasks, and uncovers their associations with creative performance.