Pith. sign in

REVIEW 3 major objections 6 minor 74 references

Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions

T0 review · 3 major / 6 minor · reviewed 2026-07-12 · grok-4.5

Pith's one-line read Hallucination is named the top OSINT-AI risk in over twenty studies, yet end-to-end rates are measured in only one system.

desk verdict Solid survey that makes the hallucination–validation gap and lifecycle imbalance usable for the field; the “only one” count is carefully scoped but still rests on a single-screener curated corpus. read the letter →

arxiv 2607.03233 v1 pith:QD4FF2LU submitted 2026-07-03 cs.CR cs.AIcs.IRcs.SI

classification cs.CRcs.AIcs.IRcs.SI
keywords open-sourceintelligenceOSINTagenticAIlargelanguagemodelshallucinationretrieval-augmentedgenerationcyberthreatevaluationbenchmarks
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Manual open-source intelligence can no longer keep up with the volume of public digital information. This survey of 74 studies argues that large language models and agentic systems—models that plan, call tools, and iterate—can help, but published demos have outrun the evaluation needed to trust them. Its central finding is a corpus-level hallucination–validation gap: more than twenty papers treat hallucination as a major reliability threat, yet only one OSINT-specific system reports an end-to-end rate, 4 percent under favourable, non-reproducible conditions; related correction results come from general question answering, not OSINT. Mapping work onto the intelligence lifecycle shows collection and analysis are well covered while verification, reporting, dissemination, and decision support are thin. The authors treat agentic AI as its own category, offer an eleven-part taxonomy, and close with a ten-point agenda and a concrete near-term stance: use these systems as co-pilots for collection and triage while analysts keep verification and decisions.

What carries the argument

The hallucination–validation gap: a corpus-level count that treats end-to-end OSINT hallucination measurement as distinct from general-domain reasoning or factual-correction scores, and uses that distinction plus an OSINT-lifecycle coverage map and an eleven-category taxonomy that separates agentic architectures from ordinary LLM prompting to drive a ten-point research agenda and the co-pilot deployment claim.

What would settle it

A second, independent OSINT-specific study that reports an end-to-end hallucination rate under open, reproducible, and preferably adversarial conditions on a public dataset would break the “only one measurement” claim and force the gap diagnosis to be rewritten.

Watch

Extended reading notes

Core claim

Across a 74-study corpus, the field names hallucination a primary reliability concern in more than twenty papers but empirically measures end-to-end hallucination on an OSINT system with OSINT data in only one case—a RAG architecture reporting 4 percent under favourable, private conditions—while agentic systems are tested only under benign inputs, no shared open OSINT-AI benchmark exists, and the OSINT lifecycle is heavily skewed toward collection and analysis rather than verification, reporting, dissemination, and decision support.

Load-bearing premise

That an expert-curated, mostly English, single-screener set of 74 studies from a limited search window is complete and unbiased enough for counts such as “only one OSINT end-to-end hallucination measurement” and the workflow-stage tallies to stand as field-level facts rather than selection artifacts.

Editorial extensions

If this is right

  • Near-term OSINT and cyber-investigation deployments should keep humans responsible for verification and decisions while models handle collection and triage.
  • New OSINT-AI papers should treat end-to-end hallucination measurement on OSINT data as a required evaluation component, not an optional extra.
  • The community needs an open, shared OSINT-AI benchmark that covers collection accuracy, cybersecurity NER, product factual accuracy, hallucination rate, analyst task completion with and without AI, and adversarial robustness.
  • Agentic systems must be re-evaluated under adversarial and contaminated inputs rather than only benign tool outputs.
  • Dark-web, multimodal, multilingual, and legally admissible collection methods become priority gaps rather than side topics.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Procurement that relies on multiple-choice cybersecurity knowledge scores will systematically overestimate readiness for open-ended threat-intelligence workflows.
  • Until adversarial filtering is built into knowledge-graph and RAG pipelines, the same grounding mechanisms that cut incidental hallucination can amplify poisoned open-source feeds.
  • The co-pilot stance implies measurable analyst LLM literacy and oversight effectiveness studies, not only architectural checkpoint diagrams.
  • A community working group for OSINT-AI benchmarks, analogous to shared NLP suites, is the institutional step most directly implied by the non-cumulative metrics pattern.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This survey reviews 74 studies on agentic and generative AI for OSINT, CTI, and cyber investigation. It contributes an 11-category taxonomy that treats agentic AI as distinct from simple LLM prompting; a corpus-level hallucination–validation gap (hallucination named in >20 studies but end-to-end OSINT hallucination measured in only one system, Allam’s RAG result of 4% under private/favourable conditions); an OSINT-lifecycle coverage map showing collection and analysis dominate while verification, reporting, dissemination, and decision support are thin; and a ten-point research agenda culminating in a human–AI co-pilot deployment recommendation. Methodology follows PRISMA-style principles with a 21-field evidence matrix on primary/methodological papers, explicit inclusion/exclusion tiers, and a substantial limitations section.

Significance. If the synthesis holds, the paper is a timely and useful survey for IEEE Communications Surveys & Tutorials: it organises a fast-moving, fragmented literature; makes evaluation design (CyberMetric vs CyberThreat-Eval) a first-class lesson rather than a scoreboard; and converts documented gaps into a concrete agenda (standardised hallucination protocols, open benchmarks, adversarial evaluation, dark-web legality, LLM literacy). Strengths include careful scoping of end-to-end OSINT hallucination versus general-domain QA correction metrics, transparent extraction fields (including a binary hallucination-measurement field), honest limitations (single screener, expert curation, preprint reliance, private datasets), and a co-pilot conclusion grounded in convergent evidence rather than architectural hype. These are genuine contributions to evaluation culture in OSINT AI.

major comments (3)
  1. [Abstract; §I-B; §III-I; §II-B; §VII] Abstract, §I-B Contributions, and §III-I: the load-bearing claim that end-to-end OSINT hallucination is measured in “only one” system is presented as a field-level finding that drives the agenda and co-pilot conclusion. Within the extracted 48-paper matrix this is transparent, but §II-B and §VII state that the corpus is expert-curated (not an exhaustive multi-database Boolean export), single-screener coded, English-dominant, and time-bounded. A negative existence claim is only as strong as coverage. Please (i) dual-code or second-review the hallucination-measurement field for all primary/methodological papers and report agreement; (ii) state consistently in Abstract/Contributions that the count is “within the reviewed corpus”; and (iii) add a short sensitivity note on how 1–2 additional OSINT end-to-end rates would affect the gap framing and Agenda Point 1.
  2. [§V-A; Table IX; Conflict of Interest] §V-A Case Study 1 elevates Palmieri et al. [4] as “the corpus’s central paper and its most complete agentic OSINT proof-of-concept,” and the same work is repeatedly the agentic reference in §III-C and Table IX. First author of the survey is also first author of [4]. Selection of flagship case studies should be justified by pre-stated criteria (architectural completeness, evaluation depth, tool coverage) independent of authorship, with an explicit author-relationship disclosure in the case-study section and/or Conflict of Interest statement so readers can assess self-preference risk. Without that, the “central agentic reference” framing is harder to defend than the broader corpus synthesis.
  3. [§II-B; Figure 1] §II-B Search Strategy: the manuscript still contains a placeholder (“the precise search dates should be substituted here if a reviewer requires an exact, replayable log”) and declines per-database hit counts because of expert curation. For a survey whose central claims are corpus-level counts and absences, please replace the placeholder with actual search window dates, list databases/repositories queried, and give at least approximate screening numbers already sketched in Figure 1 so independent auditors can bound coverage. This does not require re-running a fully automated export, but the current placeholder is not acceptable in a final version.
minor comments (6)
  1. [Table II] Table II note on 75 folder listings vs 74 unique studies (cross-listing of [19]) is clear; ensure the same cross-listing rule is applied consistently in any supplementary folder-to-theme table so totals never drift.
  2. [Abstract; §III-I; Table IV] §III-I and Table IV carefully separate Allam’s 4% OSINT hallucination rate from Verify-and-Edit EM gains; keep that distinction equally explicit in the Abstract’s second contribution sentence so casual readers do not collapse the two.
  3. [§IV; Table VI; Figures 8–9] Figures 8–9 and Table VI count paper–stage appearances (multi-counting allowed); the caption already notes this—add one sentence in §IV that primary-stage sums (e.g., 57) are not unique-paper totals to avoid misreading coverage density.
  4. [Table IV; §V; §VII] Several cited items are preprints/theses ([5], [6], etc.); §VII flags this—consider a small marker in Table IV or case-study headers (e.g., “preprint”) so evaluation maturity is visible without flipping to Limitations.
  5. [§III-A; References] Minor prose polish: occasional doubled words/phrases in the provided text (e.g., “tasks tasks,” “may may”) and incomplete trailing reference list in the source dump—proofread the camera-ready carefully.
  6. [Data Availability Statement] Data Availability points to a GitHub supplementary matrix; confirm the repo is public and stable before publication, and cite the commit or release tag if possible.

Circularity Check

1 steps flagged · score 2.0 of 10

Minor self-referential elevation of first-author Palmieri [4] as the corpus’s “central” agentic PoC and Case Study 1; main gap claims (hallucination count, workflow tallies) rest on external-corpus coding, not definitional reduction.

  1. self citation load bearing [§V-A Case Study 1; also Abstract, §I-B, §III-C, Table IX]
    "Palmieri et al. [4] is the corpus's central paper and its most complete agentic OSINT proof-of-concept. ... Palmieri thus sets the state of the art against which architectural alternatives [5] and evaluation advances [3] are measured throughout this review"

    First author Palmieri’s own prior system is elevated by the present authors as the defining/central agentic reference and Case Study 1 that “sets the state of the art.” This is self-citation used to anchor the agentic category and comparative architecture narrative. It is not fully load-bearing for the paper’s primary corpus-level findings (the “only one” hallucination measurement, workflow-stage imbalance, or co-pilot recommendation), which are grounded in the broader 74-item coding; hence only mild circularity.

full rationale

This is a systematic survey/SLR, not a first-principles derivation or predictive model with equations, fitted parameters, uniqueness theorems, or ansatzes. The load-bearing claims (hallucination measured end-to-end in only one OSINT system [6]; collection/analysis dominate while verification/reporting/decision-support are sparse; no shared open benchmark; co-pilot as near-term architecture) are synthesized from coding of a 74-item external corpus via an explicit 21-field matrix and PRISMA-style flow. They do not reduce by construction to the authors’ own inputs. The sole mild circularity pattern is self-citation load-bearing of limited scope: first-author Palmieri’s prior work [4] is repeatedly framed as “the corpus’s central paper,” “most complete agentic OSINT proof-of-concept,” and Case Study 1 that “sets the state of the art,” which is self-referential framing rather than an independent external benchmark. That elevation is not required for the gap counts or the ten-point agenda; those survive if [4] is demoted. No fitted-input-called-prediction, no uniqueness imported from authors, no ansatz smuggled via self-citation, and no renaming of a known result as a novel derivation. Per the default and hard rules, score 2 (one minor non-load-bearing self-citation) is appropriate; a higher score would manufacture circularity the text does not support.

Assumptions & free parameters 0 free parameters · 4 assumptions · 3 invented entities

As a systematic review, the paper introduces almost no fitted physical parameters; its load-bearing commitments are methodological axioms (what counts as OSINT-relevant AI, how hallucination is defined as end-to-end OSINT measurement vs adjacent QA error, OSINT lifecycle stage labels) and named analytical constructs (hallucination–validation gap, co-pilot deployment model). Free parameters are effectively absent; invented entities are organizational/synthesis constructs rather than physical objects.

assumptions (4)
  • ad hoc to paper End-to-end OSINT hallucination is defined as the proportion of fabricated or unsupported assertions in final intelligence output of an OSINT system on OSINT data, and is distinct from general-domain multi-hop QA exact-match gains (e.g., Verify-and-Edit).
    This definitional cut is what allows the claim that only one system ([6]) closes the measurement gap; stated in Abstract and §III-I.
  • domain assumption The OSINT lifecycle can be partitioned into collection, processing, enrichment, analysis, verification, reporting, dissemination, decision support, and human review for coverage counting.
    Stage map in §IV / Table VI drives the claim of structural imbalance; standard in intelligence practice but operationalized here for paper-stage tallies.
  • ad hoc to paper Expert-curated inclusion within a conceptual Scopus-style string plus forward/backward citation is sufficient to support corpus-level statements about the field.
    §II-B and §VII acknowledge non-exhaustive database export; the central gap claims depend on this representativeness assumption.
  • domain assumption Benign-only evaluation of agentic OSINT does not establish reliability under adversarial OSINT channel conditions documented by fake-CTI and misinformation studies.
    Used throughout §III-C, §III-J, and case-study synthesis to downgrade autonomous-deployment readiness.
invented entities (3)
  • Hallucination–validation gap (corpus-level construct)
    purpose: Name the mismatch between frequent acknowledgment of hallucination and rare OSINT-specific end-to-end measurement.
    Central contribution label; independent_evidence is the authors’ extraction count over their corpus, not an external benchmark suite.
  • 11-category taxonomy of agentic/generative AI for OSINT
    purpose: Organize heterogeneous literature into foundations, workflows, agents, APIs, RAG/KG, CTI, prompting, fine-tuning, evaluation, risk, dark web.
    Organizational invention of the survey; useful but not independently measured outside this paper’s coding.
  • Human–AI co-pilot near-term deployment architecture (normative synthesis)
    purpose: Recommend LLMs for collection/triage with analysts retaining verification and decision responsibility.
    Convergent recommendation from gaps, not a tested system with outcome metrics in this paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions." pith.science (2026). https://pith.science/paper/QD4FF2LU

@misc{pith2026260703233,
  author       = {Pith},
  title        = {Pith review of: Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QD4FF2LU}},
  note         = {Machine review of arXiv:2607.03233}
}
read the original abstract

The rapid growth of publicly available digital information has rendered manual open-source intelligence (OSINT) analysis insufficient for modern intelligence, cybersecurity, and cyber investigation. Large language models (LLMs) and agentic AI systems, capable of tool use, multi-step reasoning, and iterative intelligence generation, have emerged as promising solutions, yet evaluation frameworks have not kept pace with reported capabilities. This survey systematically reviews 74 studies and makes four contributions. First, it establishes agentic AI as a distinct analytical category rather than an extension of LLM prompting, organising the literature through an 11-category taxonomy covering LLM foundations, agentic architectures, retrieval-augmented generation (RAG), knowledge graphs, prompt engineering, domain adaptation, evaluation benchmarks, and risk. Second, it identifies the hallucination-validation gap as a corpus-level finding: although hallucination is recognised as a major reliability concern in over twenty studies, end-to-end hallucination is empirically measured in only one OSINT-specific RAG-based system, non-reproducible conditions, while related reasoning and factual-correction studies evaluate general-domain question answering rather than OSINT. Third, it maps existing research to the OSINT lifecycle, showing strong support for collection and analysis but limited coverage of verification, reporting, dissemination, and decision support. Fourth, it derives a ten-point research agenda addressing evaluation, benchmarking, hallucination measurement, adversarial robustness, dark-web coverage, multimodal intelligence, and governance. It concludes that a human-AI co-pilot model, where LLMs assist collection and triage while analysts retain responsibility for verification and decision-making, represents the most defensible near-term deployment architecture.

Figures

Figures reproduced from arXiv: 2607.03233 by the authors.

Figure 1
Figure 1. PRISMA-style corpus selection and classification flow. Source counts [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Number of Studies by Analytical Category [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. OSINT Relevance Rating Distribution by Analytical Tier [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Hierarchical Taxonomy Map — Agentic and Generative AI in OSINT and Cyber Investigation. The taxonomy comprises eleven categories organised [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Composite agentic OSINT reasoning loop, synthesising components described across four studies: the ReAct reasoning loop [ [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: RAG and Knowledge Graph Enhanced OSINT Architecture. This [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: Responsible Agentic OSINT Safeguard Framework. A five-layer safeguard architecture synthesised from corpus contributions. Each layer is annotated [PITH_FULL_IMAGE:figures/full_fig_p019_7.png]
Figure 8
Figure 8. Figure 8: Bubble Chart of Studies versus OSINT Workflow Stages. Bubble size reflects the number of papers addressing each stage as primary or secondary [PITH_FULL_IMAGE:figures/full_fig_p019_8.png]
Figure 9
Figure 9. Figure 9: Distribution of Studies by OSINT Workflow Stage. Bars show the count of paper-stage assignments per stage, disaggregated into primary coverage [PITH_FULL_IMAGE:figures/full_fig_p020_9.png]
Figure 10
Figure 10. Figure 10: Number of Studies by Primary Technology Type [PITH_FULL_IMAGE:figures/full_fig_p021_10.png]
Figure 11
Figure 11. Figure 11: OSINT-to-CTI and Cyber Investigation pipeline: a left-to-right flow across six stages from source environments to final output, with study identifiers [PITH_FULL_IMAGE:figures/full_fig_p025_11.png]
Figure 12
Figure 12. Figure 12: Comparative Architecture of the Four Case Studies. Each panel shows the internal architecture, tool integrations, key metrics (where reported), and [PITH_FULL_IMAGE:figures/full_fig_p029_12.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

74 extracted references · 15 linked inside Pith

  1. [4]

    A framework for embedding generative and agentic AI in open source intelligence,

    E. A. Palmieri, M. C. Ghanem, V . Sowinski-Mydlarz, and D. Dunsin, “A framework for embedding generative and agentic AI in open source intelligence,” inProceedings of the 2025 IEEE Cyber Education and Research Conference (CERC). London, UK: IEEE, 2025

  2. [1]

    Current status and security trend of OSINT,

    Y .-W. Hwang, I.-Y . Lee, H. Kim, H. Lee, and D. Kim, “Current status and security trend of OSINT,”Wireless Communications and Mobile Computing, p. 1290129, 2022

  3. [2]

    Implications of large language models for OSINT: Assessing the impact on information acquisition and analyst expertise in prompt engineering,

    J. ˇCern´y, “Implications of large language models for OSINT: Assessing the impact on information acquisition and analyst expertise in prompt engineering,” inProceedings of the 23rd European Conference on Cyber Warfare and Security (ECCWS), 2024, prague University of Economics and Business

  4. [3]

    CyberThreat-Eval: Can large language models automate real-world threat research?

    X. Chen, X. Feng, S. Chen, M. Maitre, S. Rakshit, D. Duvieilh, A. Picone, and N. Tang, “CyberThreat-Eval: Can large language models automate real-world threat research?”Transactions on Machine Learn- ing Research (TMLR), November 2025, arXiv:2603.09452; Microsoft Research / HKUST

  5. [5]

    LLM-based OSINT agent with memory, knowledge integration, tool application, and self-reflection,

    Z. Shen, Q. Wu, and K. Shen, “LLM-based OSINT agent with memory, knowledge integration, tool application, and self-reflection,” Preprint under review, 2024, tsinghua University

  6. [6]

    The impact of artificial intelligence on OSINT technologies,

    E. Allam, “The impact of artificial intelligence on OSINT technologies,” Bachelor’s Thesis, LUISS Guido Carli University, 2025, department of Business and Management; Supervisor: Prof. Gianluigi Me

  7. [7]

    A comprehensive overview of large language models (LLMs) for cyber defences: Opportunities and direc- tions,

    M. Hassanin and N. Moustafa, “A comprehensive overview of large language models (LLMs) for cyber defences: Opportunities and direc- tions,” arXiv:2405.14487v1, 2024, western Sydney University / UNSW Canberra

  8. [8]

    Review of generative AI methods in cybersecurity,

    Y . Yigit, W. J. Buchanan, M. G. Tehrani, and L. Maglaras, “Review of generative AI methods in cybersecurity,” arXiv:2403.08701v2, 2024, edinburgh Napier University

Show all 74 references
  1. [9]

    A comprehensive approach for enhancing OSINT through leveraging LLMs,

    G. Rajendran, A. Arun Kumar, P. K. Sridhar, K. K. Perumalsamy, and N. Srinivasan, “A comprehensive approach for enhancing OSINT through leveraging LLMs,”International Refereed Journal of Engineer- ing and Science (IRJES), vol. 13, no. 2, pp. 61–66, 2024

  2. [10]

    KIA: Know it all — an all-inclusive OSINT tool,

    G. S. Nagra, N. Shukla, D. Jadhav, A. Patil, and I. Khan, “KIA: Know it all — an all-inclusive OSINT tool,” in2024 International Conference on Electrical Electronics and Computing Technologies (ICEECT). IEEE, 2024. IEEE COMMUNICATIONS SURVEYS & TUTORIALS, VOL. XX, NO. X, 2026 35

  3. [11]

    Language models are few-shot learners,

    T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askellet al., “Language models are few-shot learners,” inAdvances in Neural Information Processing Systems, vol. 33. Curran Associates, Inc., 2020, pp. 1877–1901, GPT...

  4. [12]

    Chain-of-thought prompting elicits reasoning in large language models,

    J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V . Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” inAdvances in Neural Information Processing Systems, vol. 35. Curran Associates, Inc., 2022, pp. 24 824–24 8...

  5. [13]

    Artificial intelligence in data analysis for open-source investigations,

    T.-C. R ˘adoi, “Artificial intelligence in data analysis for open-source investigations,” in2023 IEEE 15th International Conference on Elec- tronics, Computers and Artificial Intelligence (ECAI). IEEE, 2023

  6. [14]

    Redefining OSINT software architecture with system-centric architecture design: A framework shaped by QAW, ADD, and ATAM,

    G. Yurtalan and S. Arslan, “Redefining OSINT software architecture with system-centric architecture design: A framework shaped by QAW, ADD, and ATAM,”IEEE Access, vol. 13, pp. 71 456–71 480, 2025

  7. [15]

    LLMs for web reconnaissance detection,

    C. Nil ˘a and V .-V . Patriciu, “LLMs for web reconnaissance detection,” Journal of Military Technology, vol. 6, no. 2, pp. 41–46, 2023

  8. [16]

    Privacy and security concerns in generative AI: A comprehensive survey,

    A. Golda, K. Mekonen, A. Pandey, A. Singh, V . Hassija, V . Chamola, and B. Sikdar, “Privacy and security concerns in generative AI: A comprehensive survey,”IEEE Access, vol. 12, pp. 48 126–48 144, 2024

  9. [17]

    Cyber- Metric: A benchmark dataset based on retrieval-augmented generation for evaluating LLMs in cybersecurity knowledge,

    N. Tihanyi, M. A. Ferrag, R. Jain, T. Bisztray, and M. Debbah, “Cyber- Metric: A benchmark dataset based on retrieval-augmented generation for evaluating LLMs in cybersecurity knowledge,” arXiv:2402.07688v2, 2024, technology Innovation Institute, UAE

  10. [18]

    Verify-and-edit: A knowledge-enhanced chain-of-thought framework,

    R. Zhao, X. Li, S. Joty, C. Qin, and L. Bing, “Verify-and-edit: A knowledge-enhanced chain-of-thought framework,” inProceedings of the 61st Annual Meeting of the Association for Computational Lin- guistics (ACL 2023), 2023, arXiv:2305.03268; NTU / Alibaba DAMO Academy

  11. [19]

    Generating fake cyber threat intelligence using transformer-based models,

    P. Ranade, A. Piplai, S. Mittal, A. Joshi, and T. Finin, “Generating fake cyber threat intelligence using transformer-based models,” in2021 IEEE International Joint Conference on Neural Networks (IJCNN). IEEE, 2021, arXiv:2102.04351v3; UMBC

  12. [20]

    GeoLocator: A location-integrated large multimodal model for inferring geo-privacy,

    Y . Yang, S. Wang, D. Li, Y . Zhang, S. Sun, and J. He, “GeoLocator: A location-integrated large multimodal model for inferring geo-privacy,” Preprint, 2024, university of Southern California

  13. [21]

    OSINT or BULLSHINT? exploring open-source intelligence tweets about the russo-ukrainian war,

    J. Niu, M. Stillman, and A. Kruspe, “OSINT or BULLSHINT? exploring open-source intelligence tweets about the russo-ukrainian war,” Work- shop Paper, 2025, munich University of Applied Sciences

  14. [22]

    Application analysis of generative artificial intelligence in the field of open source intelligence,

    L. Zhou, Y . Qin, S. Yan, G. Zhang, and L. Hu, “Application analysis of generative artificial intelligence in the field of open source intelligence,” in2024 39th Youth Academic Annual Conference of Chinese Association of Automation (YAC). IEEE, 2024, cID-encoded PDF; author li...

  15. [23]

    Development of a gamification appli- cation “Osint Maniac

    S. Siddiq and N. Qomariasih, “Development of a gamification appli- cation “Osint Maniac” to enhance open-source intelligence (OSINT) skills,” in2024 10th International Conference on Education and Tech- nology (ICET). IEEE, 2024

  16. [24]

    Exploring OSINT for modern day reconnaissance,

    A. S. Ram, A. Thalakkat, D. Udayan, A. A. Khader, H. K, and V . George, “Exploring OSINT for modern day reconnaissance,” in2023 Annual International Conference on Emerging Research Areas / International Conference on Intelligent Systems (AICERA/ICIS). IEEE, 2023

  17. [25]

    A standalone OSINT agency for stronger national secu- rity,

    D. Gauthier, “A standalone OSINT agency for stronger national secu- rity,” NSI Law and Policy Paper, National Security Institute, George Mason University, January 2025

  18. [26]

    Forecasting russian equipment losses using time series and deep learning models,

    J. Teagan, “Forecasting russian equipment losses using time series and deep learning models,” arXiv:2509.07813v1, 2025

  19. [27]

    AI-assisted OSINT/SOCMINT for safe- guarding borders: A systematic review,

    A. Karakikes and K. Kotis, “AI-assisted OSINT/SOCMINT for safe- guarding borders: A systematic review,”Information, vol. 16, no. 12, p. 1095, 2025

  20. [28]

    The not yet exploited goldmine of OSINT: Opportunities, open chal- lenges and future trends,

    J. Pastor-Galindo, P. Nespoli, F. G ´omez M´armol, and G. Mart´ınez P´erez, “The not yet exploited goldmine of OSINT: Opportunities, open chal- lenges and future trends,”IEEE Access, vol. 8, pp. 10 282–10 304, 2020

  21. [29]

    An AI agent and large language model-based approach to open source intelligence analysis,

    Y . Su, “An AI agent and large language model-based approach to open source intelligence analysis,” in2025 IEEE 8th Advanced Information Technology, Electronic and Automation Control Conference (IAEAC). IEEE, 2025

  22. [30]

    Empowering LLMs with toolkits: An open-source intelligence acquisition method,

    X. Yuan, J. Wang, H. Zhao, T. Yan, and F. Qi, “Empowering LLMs with toolkits: An open-source intelligence acquisition method,”Future Internet, vol. 16, no. 12, p. 461, 2024

  23. [31]

    LLM agent for disinformation detection based on DISARM framework,

    K. Tseng and M.-K. Shan, “LLM agent for disinformation detection based on DISARM framework,” inProceedings of the AAAI Workshop on Disinformation Detection (DEFACTIFY), 2025, national Chengchi University

  24. [32]

    OSINT clinic: Co-designing AI- augmented collaborative OSINT investigations for vulnerability assess- ment,

    A. Mukhopadhyay and K. Luther, “OSINT clinic: Co-designing AI- augmented collaborative OSINT investigations for vulnerability assess- ment,” inProceedings of the 2025 ACM CHI Conference on Human Factors in Computing Systems. ACM, 2025, pp. 1–22, virginia Tech

  25. [33]

    Auto- mated API docs generator using generative AI,

    P. Dhyani, S. Nautiyal, S. Dhyani, P. Chaudhary, and A. Negi, “Auto- mated API docs generator using generative AI,” in2024 IEEE Inter- national Students’ Conference on Electrical, Electronics and Computer Science (SCEECS). IEEE, 2024

  26. [34]

    Cybercheck: OS- INT and web vulnerability scanner,

    P. Shamunesh, S. Vinoth, and L. N. B. Srinivas, “Cybercheck: OS- INT and web vulnerability scanner,” in2023 International Conference on Electronics, Communication and Aerospace Technology (ICECAA). IEEE, 2023

  27. [35]

    OPEN EYE: An information gathering tool using OSINT framework,

    A. M. Sermakani, P. Sreejith, A. Leela Krishna, S. Lokesh Kaushik, C. V . R. Narayanan, and N. Jaya Siva Subramaniam, “OPEN EYE: An information gathering tool using OSINT framework,” in2024 Interna- tional Conference on Smart Technologies for Sustainable Development Goals (ICS...

  28. [36]

    OSINT at CT2 — AI-generated text detection: Tracing thought: Using chain-of-thought reasoning to identify the LLM behind AI-generated text,

    S. Agrahari and S. R. Singh, “OSINT at CT2 — AI-generated text detection: Tracing thought: Using chain-of-thought reasoning to identify the LLM behind AI-generated text,” inAAAI 2025 DEFACTIFY 4.0 Workshop, 2025, arXiv:2504.16913v1; IIT Guwahati

  29. [37]

    OSINT for B2B platforms,

    V . F. Pais and D. S. Ciobanu, “OSINT for B2B platforms,” in2014 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM). IEEE, 2014

  30. [38]

    Use and abuse of personal information, Part I: Design of a scalable OSINT collection engine,

    E. Rheault, M. Nerayo, J. Leonard, J. Kolenbrander, C. Henshaw, M. Boswell, and A. J. Michaels, “Use and abuse of personal information, Part I: Design of a scalable OSINT collection engine,”Journal of Cybersecurity and Privacy, vol. 4, no. 3, pp. 572–593, 2024

  31. [39]

    Open-source intelligence analysis method based on fine-tuned large models and knowledge graphs,

    Y . Su, “Open-source intelligence analysis method based on fine-tuned large models and knowledge graphs,” in2025 IEEE 8th Advanced Information Technology, Electronic and Automation Control Conference (IAEAC). IEEE, 2025

  32. [40]

    Enhancing autonomous system security and resilience with generative AI: A com- prehensive survey,

    M. Andreoni, W. T. Lunardi, G. Lawton, and P. Thakkar, “Enhancing autonomous system security and resilience with generative AI: A com- prehensive survey,”IEEE Access, vol. 12, 2024

  33. [41]

    Examine the role of generative AI in enhancing threat intelligence and cyber security measures,

    V . R. Saddi, S. K. Gopal, S. Dhanasekaran, M. S. Naruka, and A. S. Mohammed, “Examine the role of generative AI in enhancing threat intelligence and cyber security measures,” in2024 2nd International Conference on Disruptive Technologies (ICDT). IEEE, 2024, pp. 537– 542

  34. [42]

    A practical guide for OSINT investigators to combat disinformation and fake reviews driven by AI (ChatGPT),

    N. Dekens, “A practical guide for OSINT investigators to combat disinformation and fake reviews driven by AI (ChatGPT),” Industry Report, ShadowDragon, May 2023

  35. [43]

    Extraction of subjective information from large language models,

    A. Kobayashi and S. Yamaguchi, “Extraction of subjective information from large language models,” in2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC). IEEE, 2024

  36. [44]

    Political sentiment analysis on twitter using deep learning and LLM models,

    K. Mouthami, P. Naren, and R. Pranesh, “Political sentiment analysis on twitter using deep learning and LLM models,” in2025 3rd International Conference on Advancements in Electrical, Electronics, Communication, Computing and Automation (ICAECA). IEEE, 2025

  37. [45]

    Towards better cyber security consciousness: The ease and danger of OSINT tools in exposing critical infrastructure vulnerabilities,

    M. H. Pervez, N. Z. Naqvi, M. I. Ecevit, R. Creutzburg, and H. Dag, “Towards better cyber security consciousness: The ease and danger of OSINT tools in exposing critical infrastructure vulnerabilities,” in2023 8th International Conference on Computer Science and Engineering (U...

  38. [46]

    Features and implementation of DarkLens: A conceptual tool for advancing open source intelligence (OSINT) through the dark web,

    A. Amin, “Features and implementation of DarkLens: A conceptual tool for advancing open source intelligence (OSINT) through the dark web,” in2025 International Conference on Communication Technologies (ComTech). IEEE, 2025

  39. [47]

    A study on performance improvement of prompt engineering for generative AI with a large language model,

    D. Park, G.-t. An, C. Kamyod, and C. G. Kim, “A study on performance improvement of prompt engineering for generative AI with a large language model,”Journal of Web Engineering, vol. 22, no. 8, 2024

  40. [48]

    Open-source intelligence decision-making and analysis method based on large language models (LLMs),

    Q. Sun, J. Luo, Y . Tang, K. Zhang, and J. Xu, “Open-source intelligence decision-making and analysis method based on large language models (LLMs),” in2025 37th Chinese Control and Decision Conference (CCDC). IEEE, 2025

  41. [49]

    Fine-tuning large language models for sentiment classification of AI-related tweets,

    M. J. A. Riad, R. Debnath, M. R. Shuvo, F. J. Ayrin, N. Hasan, A. A. Tamanna, and P. Roy, “Fine-tuning large language models for sentiment classification of AI-related tweets,” in2024 IEEE WIE Conference on Electrical and Computer Engineering (WIECON-ECE). IEEE, 2024

  42. [50]

    ThreatModeling-LLM: Automating threat modeling using large lan- guage models for banking system,

    S. Yang, T. Wu, S. Liu, D. Nguyen, S. Jang, and A. Abuadbba, “ThreatModeling-LLM: Automating threat modeling using large lan- guage models for banking system,” arXiv:2411.17058v1, 2024, CSIRO Data61

  43. [51]

    Can generative AI models extract deeper sentiments as compared to IEEE COMMUNICATIONS SURVEYS & TUTORIALS, VOL. XX, NO. X, 2026 36 traditional deep learning algorithms?

    M. Anas, A. Saiyeda, S. S. Sohail, E. Cambria, and A. Hussain, “Can generative AI models extract deeper sentiments as compared to IEEE COMMUNICATIONS SURVEYS & TUTORIALS, VOL. XX, NO. X, 2026 36 traditional deep learning algorithms?”IEEE Intelligent Systems, vol. 39, no. 2, pp...

  44. [52]

    Do generative AI tools ensure green code? an investigative study,

    R. Sikand, R. Mehra, V . S. Sharma, V . Kaulgud, S. Podder, and A. P. Burden, “Do generative AI tools ensure green code? an investigative study,” inProceedings of the 2024 ACM Workshop on Responsible AI Engineering (RAIE). ACM, 2024

  45. [53]

    Evaluating the us- ability of LLMs in threat intelligence enrichment,

    S. Srikanth, M. Hasanuzzaman, and F. T. Meem, “Evaluating the us- ability of LLMs in threat intelligence enrichment,” arXiv:2409.15072v1, 2024, university of Guelph

  46. [54]

    Evaluation of LLM chatbots for OSINT-based cyber threat awareness,

    S. Shafee, A. Bessani, and P. M. Ferreira, “Evaluation of LLM chatbots for OSINT-based cyber threat awareness,”Expert Systems with Appli- cations, vol. 261, p. 125509, 2025, LASIGE, Universidade de Lisboa

  47. [55]

    A comprehensive overview of large language models,

    H. Naveed, A. U. Khan, S. Qiu, M. Saqib, S. Anwar, M. Usman, N. Akhtar, N. Barnes, and A. Mian, “A comprehensive overview of large language models,” arXiv:2307.06435v9, 2024, university of Western Australia

  48. [56]

    Advancements in generative AI: A comprehensive review of GANs, GPT, autoencoders, diffusion model, V AEs, and transformers,

    S. Bengesi, H. El-Sayed, M. K. Sarker, Y . Houkpati, J. Irungu, and T. Oladunni, “Advancements in generative AI: A comprehensive review of GANs, GPT, autoencoders, diffusion model, V AEs, and transformers,” IEEE Access, vol. 12, pp. 69 812–69 837, 2024

  49. [57]

    Challenges and applications of large language models,

    J. Kaddour, J. Harris, M. Mozes, H. Bradley, R. Raileanu, and R. McHardy, “Challenges and applications of large language models,” arXiv:2307.10169v1, 2023

  50. [58]

    Comparative analysis of generative AI models,

    G. Rani, J. Singh, and A. Khanna, “Comparative analysis of generative AI models,” in2023 International Conference on Advances in Artificial Intelligence and Computing Technologies (ICAICCIT). IEEE, 2023

  51. [59]

    Emergent abilities of large language models,

    J. Wei, Y . Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus, “Emergent abilities of large language models,”Transactions on Machine Learning Research, August 202...

  52. [60]

    Training compute-optimal large language models,

    J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark et al., “Training compute-optimal large language models,” inAdvances in Neural Information Processing Systems, vol. 35. Curran Associates, Inc., 202...

  53. [61]

    Large language models: Principles and practice,

    I. Trummer, “Large language models: Principles and practice,” in2024 IEEE 40th International Conference on Data Engineering (ICDE). IEEE, 2024, cornell University

  54. [62]

    Reasoning beyond limits: Advances and open problems for LLMs,

    M. A. Ferrag, N. Tihanyi, and M. Debbah, “Reasoning beyond limits: Advances and open problems for LLMs,”ICT Express, vol. 11, no. 6, pp. 1054–1096, 2025

  55. [63]

    Recent advances in generative AI and large language models: Current status, challenges, and perspec- tives,

    D. H. Hagos, S. Battle, and D. B. Rawat, “Recent advances in generative AI and large language models: Current status, challenges, and perspec- tives,”IEEE Transactions on Artificial Intelligence, vol. 5, no. 12, pp. 5873–5893, 2024

  56. [64]

    Research and application of artificial intelligence large language models based on feature enhance- ment,

    L. Zhuang, Q. Wang, L. Song, and P. Wu, “Research and application of artificial intelligence large language models based on feature enhance- ment,” in2024 4th International Conference on Consumer Electronics and Computer Engineering (ICCECE). IEEE, 2024

  57. [65]

    Training strategy based on game theory for generative AI,

    A. C. H. Chen, “Training strategy based on game theory for generative AI,” in2024 IEEE International Conference on Smart Systems for Applications in Electrical Sciences (ICSSES). IEEE, 2024

  58. [66]

    GroverGPT: A large language model with 8 billion parameters for quantum searching,

    H. Wang, P. Li, M. Chen, J. Cheng, J. Liu, and T. Chen, “GroverGPT: A large language model with 8 billion parameters for quantum searching,” arXiv:2501.00135v1, 2025

  59. [67]

    A large language model-based approach for automatically optimizing BIM,

    J. Wang, Q. Yang, and Y . Chen, “A large language model-based approach for automatically optimizing BIM,” in2024 43rd Chinese Control Conference (CCC). IEEE, 2024

  60. [68]

    Accelerating innovation with generative AI: AI-augmented digital prototyping and innovation methods,

    V . Bilgram and F. Laarmann, “Accelerating innovation with generative AI: AI-augmented digital prototyping and innovation methods,”IEEE Engineering Management Review, vol. 51, no. 2, 2023

  61. [69]

    Analysis of recom- mender system using generative artificial intelligence: A systematic literature review,

    M. O. Ayemowa, R. Ibrahim, and M. M. Khan, “Analysis of recom- mender system using generative artificial intelligence: A systematic literature review,”IEEE Access, vol. 12, pp. 87 742–87 766, 2024

  62. [70]

    Application of large language models for intent mining in goal-oriented dialogue systems,

    A. E. Shukhman, V . R. Badikov, and L. V . Legashev, “Application of large language models for intent mining in goal-oriented dialogue systems,” in2024 5th International Conference on Neural Networks and Neurotechnologies (NeuroNT). IEEE, 2024

  63. [71]

    At the dawn of generative AI era: A tutorial-cum-survey on new frontiers in 6G wireless intelligence,

    A. Celik and A. M. Eltawil, “At the dawn of generative AI era: A tutorial-cum-survey on new frontiers in 6G wireless intelligence,”IEEE Open Journal of the Communications Society, vol. 5, pp. 2433–2489, 2024

  64. [72]

    Computer vision and generative AI for yield prediction in digital agriculture,

    S. Majumder, Y . Khandelwal, and K. Sornalakshmi, “Computer vision and generative AI for yield prediction in digital agriculture,” in2024 2nd International Conference on Networking and Communications (IC- NWC). IEEE, 2024

  65. [73]

    Extracting domain models from textual requirements in the era of large language models,

    S. Arulmohan, M.-J. Meurs, and S. Mosser, “Extracting domain models from textual requirements in the era of large language models,” in2023 ACM/IEEE International Conference on Model Driven Engineering Languages and Systems Companion (MODELS-C). IEEE, 2023

  66. [74]

    Inte- grated method of deep learning and large language model in speech recognition,

    B. Guan, J. Cao, X. Wang, Z. Wang, M. Sui, and Z. Wang, “Inte- grated method of deep learning and large language model in speech recognition,” in2024 IEEE 7th International Conference on Electronic Information and Communication Technology (ICEICT). IEEE, 2024

Pith tools

Reviewed July 12, 2026 · model on record in the stance chip above.