Showing posts with label new inventions conference. Show all posts
Showing posts with label new inventions conference. Show all posts

Grounding AI Agents: literature vs. structured databases in the biopharma data stack

Sharing perspectives from the hubXchange 2026 Roundtable: “Assembling the data stack: what pharma needs from external knowledge in the age of AI Agents”

As biopharmaceutical R&D transitions toward autonomous AI agents, the industry is forced to re-examine the core substrate upon which these agents reason. A critical roundtable discussion hosted by Digital Science at the AI in Drug Discovery hubXchange 2026 addressed a foundational bottleneck: How do we balance unstructured literature with specialized structured databases to build an AI-ready data stack?

The consensus was clear: while LLMs excel at parsing the vast sea of academic literature, they frequently struggle with structured data, exposing a deep division between literature consumption and structured database integration. In this article, we recap key insights from the roundtable discussion, and propose what this means for biopharma companies moving forward. 

Unstructured literature: the promise and pitfalls of unstructured text

For many drug discovery teams, the primary value of Large Language Models (LLMs) today lies in their ability to conquer the sheer volume of academic literature. LLMs have dramatically optimized literature reviews by embedding text and measuring similar prompts across databases, removing the manual burden of reading endless papers.

However, relying strictly on LLMs introduces several acute pain points:

  • The ingestion bottleneck: While LLMs are excellent at reading text, they struggle to extract data from tables, images, and graphs. Roundtable participants highlighted that automated, validated methods to draw conclusions from tables—for example, answering, “Here are the top things you need to know in this table”—remain a critical visual gap.
  • The challenge of novelty: Experiments like Datasetpapers.com demonstrate that LLMs struggle to discover ‘unknown unknowns.’ True discovery requires explaining why a paper is groundbreaking or identifying novelty—a cognitive task where generative AI still struggles.
  • The content quality spectrum: Literature quality varies wildly. Curated papers from established publishers represent a far more reliable substrate than unvetted public repositories.
  • The missing negative results: A systemic bias exists where researchers are disincentivized from publishing negative results. In the AI age, dumping and organizing negative data must be simplified and incentivized (perhaps via shifts in H-index or S-index metrics) because machine learning models require negative data as much as positive data to build accurate classifiers.

‘Other’ databases: the cost and fragility of structured data

When moving beyond text consumption into computational biology and machine learning, biopharma relies on ‘other’ databases, such as structured sequence, mutation, chemical and clinical databases (e.g., UK Biobank, Gene PC, ADDI and PPMI for biomarker discovery validation).

These databases present a completely different set of challenges:

  • A scarcity of domain databases: There is a severe lack of structured databases in highly specific areas, such as amino acid mutations. Traditional structural models versus sequence-based models remain a continuous pain point.
  • The curation and funding crisis: Unlike the commercial publisher ecosystem, specialized public databases suffer from a chronic lack of ongoing funding. Representing this data, maintaining standards, and keeping it up to date is highly expensive. While large pharmaceutical corporations can absorb these costs, smaller biotech startups are effectively locked out.
  • The funding shift: Government and national lab funding is increasingly shifting away from pharmaceutical R&D toward materials and closed-loop discovery, leaving national standards (such as those from NIST) updated frequently but lacking deep biological curation resources.
  • Fouled public repositories: Without consistent curation and funding, public repositories easily become fouled with incorrect or mislabeled data. Because machine learning models are incredibly picky with data quality, many biopharma organizations are moving toward training models strictly with internal, highly-vetted data or avoiding public repositories altogether.

Harmonization: bridging the divide

The ultimate breakdown in the biopharma data stack occurs when attempting to harmonize unstructured literature findings with structured external databases and internal R&D data.

To make external data truly AI-ready, members of the roundtable discussed addressing several key operational requirements for their specific use cases and data needs:

  1. Defining clean data extraction standards: Clean data must go deeper than prose. In assay standards, data in the ‘methods’ section must be pristine and explicitly represent both positive and negative results.
  2. The metadata connection: Databases that rely on vague terms like ‘sample’ are highly problematic. Every raw read must be rigorously connected to downstream metadata, including omics data, patient profiles and study protocols.
  3. The missing donor ledger: A glaring gap in the current data stack is the lack of an interconnected, harmonized database of human blood and tissue donor identifiers.
  4. Re-identification risks: As AI agents cross-reference and harmonize independent datasets (e.g., matching blood donor IDs across omics databases), the risk of patient re-identification escalates, introducing complex legal and compliance hurdles.

Governance: build vs. buy and the role of knowledge graphs

Because external data sources are frequently unstructured, inconsistent, or lack verified provenance, biopharma organizations generally refuse to use external data for GxP-level decision-making and reporting. Instead, its use is confined to secondary exploratory reporting.

This reality forces organizations to ask: Is it worth purchasing external data and investing massive effort to prove its lineage, or is it more efficient to build it internally?

To navigate this landscape, organizations are leveraging two core architectural strategies:

  • Knowledge graphs: Tools like metaphactory, a Digital Science solution, are critical to building semantic ontologies, mapping disparate terminologies, and pointing autonomous AI agents directly to trusted data sources. They handle licensing, rights management, and legal boundaries.
  • Provenance machines: Tools like metaphactory’s metis act as black-box provenance machines, ensuring that every claim, entity, and relationship extracted by an agent is traceable back to its origin.

Systemic collaboration is needed

Ultimately, biopharma cannot rely on literature or databases in isolation. The future of AI-driven drug discovery depends on a composable substrate where unstructured literature discoveries are programmatically harmonized with highly structured external and internal databases. Building and maintaining this substrate is not a task for any single organization—it requires deep, systemic collaboration between biopharma, biotech, academic publishers, and AI tool developers to ensure the data powering tomorrow’s agents is reliable, traceable, and GxP-compliant.

Building an AI-ready data stack is a team effort. Talk to Digital Science about how metaphactory and metis can help.

The post Grounding AI Agents: literature vs. structured databases in the biopharma data stack appeared first on Digital Science.



from Digital Science https://ift.tt/P9gblqz

The next era of AI in drug discovery

As AI moves from summarizing papers to generating scientific claims, biopharma faces a hard question: how do you trust an answer with no way to trace it back to the truth? Digital Science’s Mark Hahnel unpacks the shift at hubXchange 2026.

Keynotes & insights from AI in Drug Discovery hubXchange 2026

On September 9, 2026, leaders across the biopharmaceutical and artificial intelligence sectors gathered in San Francisco for the AI in Drug Discovery hubXchange. Digital Science’s Mark Hahnel delivered a keynote address examining how AI is altering scientific inquiry and biopharma R&D. Moving beyond basic task automation, Hahnel articulated a future centered on data provenance, agentic workflows and the assembly of an industry-wide data substrate.

Read on for a recap and key takeaways from Hahnel’s keynote.

Moving beyond ‘AI as a Tool’ to agentic workflows

Scientific research has entered the ‘4th paradigm’, where massive volumes of data must be readily available and actionable. However, the current velocity of AI adoption is accelerating at a rate that introduces friction into traditional organizational workflows. Rapid theoretical milestones highlight this momentum, such as OpenAI publishing proofs for the complex Navier-Stokes fluid mechanics equations on Twitter. The tension between independent mathematicians leveraging Codex to tackle similar challenges, and claims of unethical scooping, demonstrates the need for tools that protect IP while supporting research advancement.

Large corporations are embracing AI, transitioning from basic productivity aids to deploying AI for generating new scientific knowledge and discovering novel drugs. Room consensus at hubXchange aligned with statistics indicating over 50% enterprise adoption across major organizations. The biopharma industry has officially progressed past using isolated AI tools and entered the next phase: managing siloed data alongside specialized domain models and autonomous agentic workflows.

Offline & local models: safeguarding intellectual property

As enterprise adoption deepens, maintaining strict data security and protecting early-stage IP remain paramount. To prevent proprietary research from being inadvertently exposed or ‘scooped’ through public cloud chat windows, biopharma companies are prioritizing local models and offline data architectures.

Dedicated tools like Digital Science’s Papers AI address this demand by keeping enterprise data, local models, and analytical routines fully offline, ensuring researchers can leverage modern AI capabilities without compromising security.

The Provenance Crisis & Data Trust

As generative systems output claims at scale, biopharmaceutical organizations face a foundational challenge: Where did this data originate, and how was this specific claim validated? Grounding claims in verifiable truth is critical because core human facts reside outside the latent weights of Large Language Models (LLMs).

To achieve full traceability, organizations must back up every statement and derived insight. Hahnel highlighted the framework detailed in Digital Science’s FAIR data playbook for Pharma white paper as an essential roadmap for establishing structured, trustworthy data environments. Regulatory compliance necessitates adhering to the FDA + EMA Guiding Principles of Good AI Practice in Drug Development, which require tracking the explicit source, raw underlying data, and precise timestamps for all AI-assisted findings.

Constructing & deconstructing papers for machines

Building robust drug discovery models requires looking beyond high-level literature summaries. While cheap and accessible methods exist—such as using models like Claude to ingest titles and abstracts from PubMed—true drug discovery demands deep full-text extraction. Full text is essential to extract vital scientific nuances, including detailed methods, experimental edge cases, figures, and direct scientific contradictions.

Navigating this domain requires working within a fragmented publisher landscape, where the top 5% of publishers account for approximately 61% of all scientific publications. Existing pharma licensing agreements provide a pathway to deconstruct and reconstruct scientific papers into machine-ready structures, enhancing internal proprietary data.

Data must be structured once across workflows so it can be continuously reused rather than repeatedly extracted. Maintaining these comprehensive global databases requires continuous operational maintenance; for example, maintaining Dimensions‘ global patent database requires a workforce actively liaising with patent offices worldwide to correct inaccuracies and guarantee precision.

Four core techniques for structured data extraction

To derive locally verifiable statements and establish end-to-end data provenance, four primary computational techniques are being actively deployed:

  • Mapping: Leveraging LLM-driven ontology mapping to harmonize disparate scientific terminologies across domains.
  • Graphs: Building dynamic knowledge graphs that represent evolving biological relationships and entities.
  • Triage: Implementing just-in-time triage and filtering to parse incoming streams of scientific literature efficiently.
  • Extractors: Deploying agentic extractors designed to pull out claims, experimental methods, biological entities, and explicit relationships directly from full text.

These techniques allow organizations to extract claims and harmonize them so they are composable with internal proprietary extensions, Electronic Lab Notebooks (ELN), and existing R&D workflows.

Conclusion: the substrate is the work

The primary takeaway from Mark Hahnel’s presentation is clear: while foundational models and agent frameworks will continuously improve, the ultimate value lies in the data substrate beneath them. Grounded scholarly inference requires generating answers built upon verified external data seamlessly combined with internal enterprise assets.

Building this substrate requires deep collaboration across biopharma, biotech, academic publishers, and AI tooling providers. Models and agent frameworks will continue to evolve, but establishing the underlying, composable data substrate is the foundational work that the entire field must build together.

Ready to build the data substrate your AI strategy depends on? Explore how Digital Science’s enterprise solutions help biopharma organizations turn siloed data into trusted, structured, AI-ready assets.

The post The next era of AI in drug discovery appeared first on Digital Science.



from Digital Science https://ift.tt/Zg0Jmiy

In Life Sciences, data integrity is non-negotiable

In the age of AI, research intelligence has to begin with trusted data. The Inside Our Data series explores the foundational data infrastructure that makes trust possible. 

The cost of an answer no one can explain

Enterprises rely on research intelligence to set and meet strategic objectives; intelligence is what makes an enterprise competitive. Research intelligence, as a concept, isn’t new. What is new is the enterprise’s necessary reliance on huge volumes of data—and on AI-driven insights and analysis which take that data as truth. 

Enterprises are moving fast to embed AI into research, analytics, and strategic planning. Far from the zeitgeisty pilots which were largely based on frontier model usage, AI-driven intelligence now comprises foundational infrastructure that informs how decisions are made. 

But for the advances and benefits this technology has already brought about, it has also brought risk. Research and AI-driven intelligence are only as trustworthy and defensible as the data they use, and not all data can withstand the necessary scrutiny of independent review.

This can pose an existential threat to research enterprises operating in regulated industries such as Life Sciences. As enterprises continue to evolve and embrace powerful new technologies, it’s more important than ever that their underlying data can stand up to audit. 

In this article, we’ll look at what happens when enterprises lack data integrity, what “trustworthy” data actually means in Life Sciences, and how the right infrastructure can fortify enterprise data in the world of AI.

What happens when the data underneath a scientific conclusion can’t be checked?

In 2020, two COVID-19 studies were published—one in the Lancet and one in the New England Journal of Medicine—using data from Surgisphere, a little-known analytics firm. 

But soon, there was a problem: Surgisphere refused to release its data for an independent audit. Both studies were retracted nine days apart.

The Lancet study claimed hydroxychloroquine increased mortality risk in COVID-19 patients—a finding that prompted the WHO to briefly pause a hydroxychloroquine arm of its global Solidarity trial before the retraction. The NEJM study was also retracted, but kept being cited long after: a Journal of the American Medical Association Internal Medicine analysis found 652 verified citations, with more than half of them occurring at least three months after the retractions took place. 

This incident illustrates what is at risk when data can’t be audited. We don’t know why Surgisphere wouldn’t release the data. Maybe it was all fabricated, maybe it wasn’t. There’s no way to know. But it doesn’t really matter. Data that cannot be audited is contagious. Bad data doesn’t stay where it started. It moves into papers, then models. Bad data has always been contagious. AI gives it a much higher reproduction rate.

It’s likely that the initial retractions were costly and frustrating for the firms who carried out the studies. But this incident also contributed to a wave of inaccuracy in critical research areas. Data that can’t be audited can halt clinical trials, knock percentage points off a stock valuation, cause lasting reputational damage, and most seriously, negatively impact the lives of real people. 

What “trusted data” really means

When it comes to defining what makes good data, enterprises aren’t starting from square one—Life Sciences has already formalized what “trustworthy” data means.

Regulators have relied on ALCOA—Attributable, Legible, Contemporaneous, Original, and Accurate—since the 1990s to assess data integrity in clinical and manufacturing contexts. More recently, guidance from bodies like the Medicines and Healthcare products Regulatory Agency and the World Health Organization extended this into ALCOA+, adding four further requirements: data should also be Complete, Consistent, Enduring, and Available. It’s a checklist built around whether a record can be verified after the fact, and it’s still required for trustworthy data today.

The related framework, FAIR—Findable, Accessible, Interoperable, and Reusable—addresses a different but equally consequential point of failure. Introduced in 2016, FAIR has had a substantial impact on how research-generating organizations think about their data: not just whether it exists, but whether it can be found, retrieved under clear terms, and reused with confidence in its provenance. FAIR doesn’t require data to be open to all—a dataset behind a paywall or access agreement can still be fully FAIR-compliant. However, in order to satisfy the requirements of the framework, it needs to be made available, in a FAIR manner, to the people who would make assertions on that data. 

The Surgisphere retractions occurred because the underlying datasets were non-compliant with these frameworks. The data wasn’t available or accessible for audit, which meant we also couldn’t know if it exemplified the necessary integrity that made it suitable for use in research.

These frameworks comprise a non-negotiable baseline for Life Sciences research enterprises, but there remain grey areas which can have unintended effects on the quality of research datasets. For example, an open dataset which is seemingly FAIR and ALCOA+-compliant could be skewed toward whichever countries or funders proactively volunteer their data. In a regulated environment, this isn’t enough; passing an ALCOA+ or FAIR checklist doesn’t tell you whether a dataset is representative—curation, applied on top of these frameworks, can correct for that skew. 

Data infrastructure designed to accommodate investigation

Data curation refers to the ongoing process of ensuring that data is complete and representative—a process that requires human judgment and relationships to execute. This is a foundational tenet of the datasets which comprise Dimensions by Digital Science, one of the world’s largest research and funding data repositories.

Dimensions was built around the idea that research intelligence is only useful if it can be traced across the full lifecycle it describes—not only publications, but the funding, trials, patents, and policy activity that surround them. Dimensions datasets span six linked content types: more than 165 million publications, 8.2 million grants, 74 million research datasets, 180 million patents, 976,000 clinical trials, and 2.5 million policy documents, all cross-referenced with the others. 

A model surfaces a promising area of research. Don’t just take the answer. Ask:

  • Who funded it?
  • Which researchers produced it?
  • What publications followed?
  • What datasets underpin them?
  • Were patents filed?
  • Did it progress into clinical trials?
  • Did it influence policy?

This structure enables attributability and originality under ALCOA+: publication records are enriched through full-text indexing and linked back to direct publisher partnerships, Crossref, PubMed, and other authoritative sources. The grant data comes from more than 700 funders worldwide, sourced by data experts directly from funder organizations wherever possible. Clinical trial records are pulled directly from official registries spanning every major region, so status, sponsors, and outcomes reflect the authoritative record rather than a secondhand summary. Patent data is provided by IFI Claims, curated and normalized by Digital Science teams.

Every record carries a persistent identifier and a link back to its original source, so a grant, publication, or patent is Findable and its provenance is never in question. Records are Accessible under clear, documented terms—whether that’s open data or a governed connection through a licensed platform, so users always know what they’re looking at and where it came from. And the cross-referencing between content types is what makes the data Interoperable and Reusable in practice: a grant can be traced through to the publications it funded, the datasets and patents those publications generated, and the clinical trials or policy documents that followed. Research across more than 100 countries and every major discipline reduces the blind spots that come from a literature-only view or a single-region dataset. This is data that can be audited—and that enterprises can trust to drive the decisions they make. 

The final step to unlocking truly powerful and trustworthy intelligence is ensuring this data infrastructure is in sync with enterprise-specific ontologies. Pairing trusted data with semantic definitions lays the groundwork for life sciences enterprises to more safely rely on AI-driven research intelligence in the years to come.  

Building trusted foundations for future AI implementations

The Surgisphere debacle exemplifies the failures that AI-assisted workflows now risk automating at scale: fluent, confident outputs based on data that doesn’t meet industry standards. Today, AI-assisted workflows are increasingly embedded in how R&D and Medical Affairs teams triage literature, surface signals, and make decisions. The efficiencies and insights to be gained from this technology are unprecedented, but this also raises the stakes: an AI working from ungoverned data doesn’t just produce a bad answer, it can introduce existential risk. 

The best way to guard against such a failure—and set your company up for long-term success—is to take a two-pronged approach, pairing an enterprise-specific semantic layer, such as a knowledge graph, with data that is FAIR and ALCOA+-compliant. Digital Science offers knowledge graph infrastructure designed to grow with an enterprise via its proprietary technology, metaphacts. 

By defining a semantic layer, an enterprise sets the scope for the data that AI is able to access and defines the logical relations between defined entities. This means the model can only interpret the data it is given access to in the context of an approved series of rules. This mitigates the risk of hallucinations or logical failures, and makes it simple for auditors to interrogate the pathways that led to a certain output. This is how to ensure trusted data is treated predictably by trusted models.

The result is accurate intelligence with an in-built audit trail that enterprises can trust to stand up to independent audit. 

The bar for data integrity will keep rising

As AI becomes more embedded in R&D and Medical Affairs decision-making, so too will audits by regulators and internal stakeholders. Trusted intelligence starts with trusted data, and trusted data is best used in sync with foundational enterprise infrastructure.

Digital Science provides one of the world’s broadest collections of connected research intelligence—combining Dimensions, Altmetric, and IFI Claims to help enterprise organizations support analytics, strategic decision-making, innovation, and AI workflows that can be explained, audited, and defended. 

The post In Life Sciences, data integrity is non-negotiable appeared first on Digital Science.



from Digital Science https://ift.tt/QRGXa1F

A double-edged sword: the growing complexity of Medical Affairs publication performance data

The variety of channels and audiences that define scientific communications reach and engagement is growing. In turn, Medical Affairs teams face diversifying data sources and tools to assess publication performance. 

Compass Points: The Future of Medical Affairs is a series exploring the strategic challenges facing medical affairs teams in today’s communication landscape—and the tools that will help them get it right.

Even the most groundbreaking data cannot change clinical practice if never translated into action. As such, a fundamental purpose of scientific communications is to inform and educate on this new data, what it means, and how it can impact the real world. The challenge is how to do this effectively across multiple regions, channels, and audiences, and how to track success (or failure).

As the complexity of scientific communication scales, Medical Affairs teams rely on an expanding library of data sources and tools to analyze the performance of scientific communications tactics. Quantifying asset performance and impact directly informs strategy, and in turn, informs publication planning. We see that this feedback loop propagates the outcomes of tactical and strategic decision-making, whether these outcomes were desirable or undesirable.

Publication planning and performance feedback loop

The growing availability of data sources and tools used to define publication performance is a double-edged sword: capabilities increase, but so, too, does workload. Assessments performed in different settings, at different time points, with non-standardized queries may create inconsistency in those outputs contributing to strategic decisions about publications. The value of scientific communications can be efficiently captured by measurement tools, such as Compass by Dimensions, characterized by integrated sources, standardized data, and intuitive performance benchmarking.

Limitations become visible when publication performance data sources and reporting tools are siloed.

As part of Medical Affairs scientific communications planning and evaluation, asset performance directly informs publications strategy. A growing variety of data sources and analysis tools are now available. These help determine publications’ reach and engagement, and by extension, their impact.

Citation tracking tools hold continued relevance. In what may represent a highly manual process, pertinent altmetrics must first be defined, then followed over time. Social media listening offers publication performance insights from an altogether different channel. To assess proprietary (or competitor) abstracts, posters, and podium presentations, congress trackers of varying complexity are commercially available or developed in-house. Whether for conferences, publishers, or individual journals, both the type and availability of performance metrics vary widely. 

These examples are not comprehensive. As their variety suggests, publication performance data sources and reporting tools are often functionally siloed from one another. They must be evaluated in turn, and the readouts integrated, to generate a comprehensive snapshot. 

Being inherently decoupled, it follows that the data sources and reporting tools illustrated here will lack technical platform interoperability. Plainly stated, they don’t communicate. As such, they are limited in their ability to provide integrated readouts and a contextual story of scientific communications asset performance.

What does this mean for user experience and workload?

Across life sciences industries, the size, structure, and distribution of Medical Affairs and publications teams differ significantly. Scientific communications strategy may be defined within the Medical Affairs functional area alone, or within a cross-functional center of excellence or integrated evidence planning team.

Where data and reporting tools are managed by a group of colleagues, only by investing time and aligning their efforts can these contributors integrate findings into a cohesive performance narrative. If such coordinating and reporting activities are repeated on a monthly basis, for example, we begin to grasp the many people-hours required. In the present era of remote work, it’s likely that these team members do not work in the same physical space, or even the same time zone. Creating the impact story requires continuous touchpoints, further decreasing efficiency.

It is important to highlight this concept of the scientific communications impact story, as creating it is just one step in the process. Another key aspect is telling that impact story effectively to leadership and other key stakeholders. How are the publication performance data contextualized? What reporting content can decision-makers expect to see, and reliably? 

A holistic scientific communications performance overview, delivered on-schedule with consistent format, takes significant time and effort, whether the overview’s creator is a team or a single contributor.

In the case of a single contributor such as the publications manager or director, this colleague is solely responsible for the time-consuming, repetitive work of integrating increasingly complex data sources and tools. Expertise more impactfully invested in key project management and strategic activities is instead diverted to data analysis. The workload risks overwhelming that colleague.

Whether in (bio)pharma, biotech, or medtech organizations, this situation’s impact may be more acutely felt in publications teams serving multiple disease or product areas. In a further example, its impact is visible in small- and medium-sized life sciences companies, where publications colleagues may “wear other hats,” having broader role descriptions or functional responsibilities.

When publication performance insights are integrated from diverse sources, how does this influence their perception?

Building on this, publications teams are facing operational environments in which scientific communications performance assessment and reporting processes become overwhelming.

While these may be subject to formalized standard operating procedures, it’s more likely that practices fluctuate over time: team structures change, or publication types evolve. Inherent knowledge informs the work of integrating performance data from diverse sources, often depending on personal best practices. Processes become opaque, and as the risks of missing relevant data and of differing interpretations increase, reporting inconsistencies emerge.

Whether monitoring owned or competitor assets, publication performance reporting serves myriad purposes. These range from publication impact measurement, to downstream budget and strategy planning, to competitive intelligence. Performance reporting is meant to describe impact and value.

Should the integrated insights appear inconsistent, this perception affects stakeholders. It reflects negatively on the work and reputation of the publications or scientific communications team, the Medical Affairs team, or the integrated evidence planning team. Cross-functional partners or leadership may perceive the accumulated insights as unreliable, or even non-actionable. Over time, this hinders effective business decision-making, perceived department value, trust, and even individual working relationships.

Data integration workarounds that utilize generative artificial intelligence lack fidelity.

In the last three years, multimodal generative artificial intelligence (genAI) technologies have gained significant traction as data integrators. Their ability to instantaneously compare inputs, summarize findings, and create personalized outputs feels reassuring. With remarkable efficiency improvements, a single user can develop polished, on-brand content and dashboards in minutes.

GenAI technologies may represent a tempting solution to the challenge of publication performance data collected from such disparate sources and tools. This is especially true for life sciences organizations holding enterprise agreements that facilitate company-managed access to these technologies.

It is critical to balance the benefits of improved efficiency against the limitations of utilizing genAI as a process workaround to analyze and integrate publication performance data. Due to these technologies’ very design, they are neither able to consistently benchmark nor to track target performance metrics over time. As such, assessments remain snapshots that must be repeated according to stakeholders’ reporting requirements.

Hallucination and sycophantic responses are known challenges with the use of genAI. Outputs with publication performance data integration as their goal may be incomplete, factually incorrect, or biased. A genAI-grounded process still relies on the user to identify and supply trusted data sources. If pertinent metrics are missing, genAI-directed data integration processes cannot account for them. Alternatively, depending on how the user prompts the model, they may have the undesirable experience of hallucinated metrics or outputs.

The use of genAI to speed up integration of disparate, disconnected data sources should not come at the cost of insight fidelity. Rather, when artificial intelligence capabilities are paired with data analytics, reliable analyses require standardized, consistent data feeds from curated sources. When a publications team builds such analytics de novo, both the data sources and analytics outputs take time to verify and to trust.

Standardization and repeatability are key to successful publication performance assessment.

Capturing the value of scientific communications should not be held back by the repetitive work of reconciling disparate data sources. Nor should strategy-defining insights depend on workarounds, themselves subject to technical limitations. As well, it is worthwhile to consider the accumulated inefficiencies that these activities create for publications managers and teams.

Measurement tools that integrate data sources by their design unlock the power of user-defined search and tracking parameters. Meaningful insights are uncovered when these parameters are standardized and repeatable, tracking publication performance with consistency over time. When unique, Medical Affairs-relevant data sources come already embedded, it streamlines the work of uncovering scientific communications reach, engagement, and impact. This diversity of data is no longer an obstacle.

Compass by Dimensions captures these capabilities. Built on more than a decade of Dimensions and Altmetric data trusted by industry, it is designed to help overcome the challenge of data diversity. Compass combines publication and altmetrics into a single collaborative workflow, reducing inefficiencies, saving time, and simplifying how publications professionals and Medical Affairs teams benchmark, track, and manage publication impact and reach.  

Compass by Dimensions is developed by Digital Science, an AI-focused technology company that transforms fragmented data into unified knowledge assets, leveraging AI and Knowledge Graphs to deliver structured, actionable intelligence for high-value discovery and innovation. By combining unparalleled data depth and breadth with enterprise-ready AI technology, we help leaders confidently accelerate product life cycles and secure a decisive market lead.

The post A double-edged sword: the growing complexity of Medical Affairs publication performance data appeared first on Digital Science.



from Digital Science https://ift.tt/tuWEXl6

Show your sources: building verifiable, citable AI agents with MCP

Model Context Protocol (MCP) is an open standard which connects LLMs with external systems. We discuss how new Dimensions and Altmetric MCPs ground LLMs in structured data, generating verified, citable results that research teams can trust.

The “Knowledge Gap” in AI

You can’t create cutting-edge research from stale, outdated information. And yet research teams are trying and failing to derive insights from standard Large Language Models (LLMs) trained on data that is months or even years old. A significant problem in research contexts where new information is constantly released, and hundreds – if not thousands – of new publications and reports are published daily. 

AI tools can take over time-consuming research tasks like competitor tracking, but connecting them to trusted data can take more custom coding and login/security setup than most teams have the resources for. When research teams lack the necessary technical abilities to hard-code, they are alienated from the successes of AI Research Integration. One in three R&D-focused enterprises say understanding and implementing AI tools is one of their biggest challenges.

With the Model Context Protocol (MCP), no team needs to miss out. In this article, we explain what an MCP is, why it’s governance-friendly, and explore how MCPs elevate strategy and innovation with Dimensions and Altmetric MCPs. 

What is Model Context Protocol (MCP)?

Large Language Models (LLMs) promise enormous potential. But the potential of these models has been stunted. This is because the data AI interacts with is siloed and trapped behind legacy systems. AI models are thus forced to use outdated training data to make their decisions, and can hallucinate when asked to reason over current scientific literature.

For example, a researcher exploring lung cancer treatments may ask an LLM to identify, based on all available scientific literature and past oncology trials, the toxicity risks of a promising drug. The LLM outputs an authoritative and neatly argued “green light” for the use of this drug, for this new application. What this researcher doesn’t realize is that not only has the LLM missed an influential dataset published last week, which evidenced harmful effects for patients with specific comorbidities, but also hallucinated a single decimal point in a critical dosage threshold. 

In a study by NVIDIA, 59% of respondents from pharmaceutical and biotech companies cited drug discovery and development among their top AI use cases.

Previously, if you wanted to integrate AI into an external service, this required lengthy custom implementations. Now, the Model Context Protocol (MCP) provides a standardized, plug-and-play protocol for AI applications to connect to external data sources in a structured, reliable, and permissioned way – similar to how a USB-C port allows external devices to connect to computers. This means autonomous agents (like Claude or ChatGPT) can query relevant data using natural language instead of complex code. 

With MCP, you can connect AI assistants like Claude, Cursor, VS Code Copilot and ChatGPT with external data sources and tools. This might include productivity tools (like Slack), development tools (like GitHub) and data and file systems (like Google Drive). 

Consider an engineering team trying to develop a new kind of lightweight battery for electric cars. Previously, when using LLMs to challenge and improve their prototypes, they would copy and paste their research into the chat window each session. By connecting AI to their data sources and tools via MCP, copying and pasting their research became unnecessary. Not only does the AI automatically query the relevant research mid-conversation, but it can also surface relevant context from a years-old study, buried in the organization’s Google Drive. This connection sparks the game-changing insight. 

MCP: the data flow your IT team will actually approve

Teams who primarily work via cloud services will be – perhaps acutely – familiar with the lengthy AI governance for research required for new SaaS tools. In these kinds of environments, AI governance requires the continuous real-time monitoring and risk assessment of every team member’s AI adoption. In a fast-paced research environment, this is a time- and resource-intensive undertaking. According to a 2025 AI-Ready Governance Report, organizations reported a 37% jump year-on-year in time spent managing AI risk, and 98% of surveyed organizations had made plans to increase their governance budgets.

Because MCP facilitates a one-way flow of trusted data into your existing tools, it’s governance-friendly. External products provide the data, but what is done with that data is entirely private to the user. Only users can make calls to the data, meaning that applications cannot see researchers’ prompts, enterprise’s internal data, or how AI is processing the information. This includes no “AI-to-AI” linkage, meaning that there is no need to be concerned that an enterprise’s internal data is leaked to MCP providers’ models. This is another reason why MCP integrations typically bypass the lengthy governance reviews required for new SaaS tools.

Discover Digital Science’s Dimensions & Altmetric MCPs

To maximize the potential of this innovation, Digital Science has designed Dimensions MCPs (Semantic Search and Analytics MCP) and Altmetric MCP to cater to the specific needs of research and data teams and provide verified, citable results so that users can discover more. 

Dimensions Semantic Search MCP translates plain-language questions into precise queries to run across one of the world’s most comprehensive databases, returning answers with full provenance. Through ontological concept resolution, the ontologically aware query, which drives the MCP, searches for concepts and corresponding synonyms. 

The Dimensions Analytics MCP provides a connected AI with a map of the research landscape, providing the links between relevant people, funding and organizations.

Altmetric shows where research is actually being read, shared, and acted on – in news media, policy documents, patents, and online conversations. Its MCP brings that attention data into your AI workflows so you can surface and report on societal impact at scale.

Key Benefits of Dimensions & Altmetric MCPs

Access to data from more than 430 million interconnected records, spanning publications, grants, patents, clinical trials, datasets, and policy documents, is just the beginning of what Dimensions and Altmetric offers research teams.

Despite the volume and complexity of the linked research data indexed, integrating Dimensions and Altmetric MCPs into a SaaS infrastructure requires neither complex code nor complex authentication. Because MCPs provide a standardization layer, users can move from setup to insights in minutes, not days, eliminating the need for heavy engineering resources.

The Altmetric MCP and two Dimensions MCPs are each designed to meet a specific need, creating a comprehensive research intelligence stack.  

Advanced content search with Dimensions Semantic Search MCP

  • Ontological concept resolution: Drug names, diseases, compounds are mapped to structured IDs across 40+ domains.
  • Co-occurrence discovery: Surface which drugs or compounds appear most with a given disease.
  • Multi-source search: Data from more than 430 million combined publications, patents, grants and clinical trials is considered in every query. 
  • Combined filters: Blend concept search with proximity, date, and author constraints.

Mapping the research ecosystem with Dimensions Analytics MCP

  • 430M+ linked records: Connect the dots between publications, grants, patents, clinical trials, datasets, policy documents, and even researchers and organizations.
  • Rich metadata: Find the most relevant data with robust metadata on research content records.
  • Natural Language & DSL: Query in plain English or use the full Dimensions Search Language.

Measuring impact with Altmetric MCP 

  • Attention Tracking: Monitor news, policy, social media, and patent references to research outputs in real time.
  • Filtering: Sort institutional research outputs by author, journal, and publication date.
  • Custom Analysis: Organize data by impact to your work – from publication-level details to aggregated metrics by therapeutic area, asset, key opinion leaders (KOL) impact, or company-level performance.

Dimensions Semantic Search MCP helps AI surface the content and evidence researchers need. It uses semantic technology to help teams query relevant material in publications, grants, patents, and clinical trials with just a key phrase or concept.

Dimensions Analytics MCP provides a linked view of the research ecosystem. It informs your AI workflow with research context from 430+ million linked publications, funding, researcher and organization profiles.

Altmetric MCP identifies domain experts and reveals the real-world impact of research and products via news, policy, social media, clinical guidelines, and patent monitoring.

Use Cases: From Discovery to Strategy

Here are some use case examples that demonstrate existing pain points teams face today, and how MCP, and particularly Dimensions and Altmetric MCPs, can help to solve their challenges.

Target Identification with Dimensions Semantic Search MCP

The problem: Researchers want to screen gene targets for a rare hereditary disease causing facial dysmorphism.

The strategy: Via the Dimensions Semantic Search MCP, the researchers are able to search in a harmonized way, using resolved concepts, across information from clinical trials, research publications, and patent data around the world and receive verified, citable results. The team is then able to visualize and generate reports with AI tools like GitHub and Microsoft Copilot, which can be integrated with Dimensions Semantic Search MCP. 

Cross-entity intelligence with Dimensions Analytics MCP

The problem: A biopharmaceutical company needs to find key opinion leaders (KOLs) and top researchers on a rare autoimmune condition for a steering committee.   

The strategy: Dimensions Analytics MCP can help the team identify key opinion leaders (KOLs) and accelerate strategic collaborations by mapping top researchers directly to their full scientific footprint. The Dimensions Analytics MCP connects you with deep, linked data on global publications, clinical trials, grant funding, intellectual property (IP), real-world impact metrics, and collaborative research networks. From broad therapeutic areas down to topics as granular as CRISPR-Cas12a off-target cleavage mechanics or AAV9 capsid engineering for crossing the blood-brain barrier in ALS, Dimensions Analytics MCP connects you to top researchers in any niche.

Research impact with Altmetric MCP 

The problem: A green chemistry firm publishes a landmark paper on a discovery they’ve made in catalyst technology, and wants to monitor the attention their breakthrough receives. 

The strategy: By linking an internal agent to the Altmetric MCP, the company can run automated reports of references to their research in real time.

Conclusion: Future-Proof Your AI Roadmap

At Digital Science, we’ve curated a range of products which help researchers push the boundary of discovery. And thanks to natural language discovery, the next big breakthrough in a research project could be unlocked via a simple prompt. In keeping with the Digital Science mission to democratize knowledge, the Dimensions Semantic Search MCP, Analytics MCP and Altmetric MCP are designed to make the research process as accessible and intuitive as possible, empowering non-technical team members to explore data from millions of publications and other research records via the built-in conversational interface or via the user’s agentic AI, linked via MCP. 

If you are interested in learning more about what AI research integration can achieve, contact us to see firsthand how the Dimensions and Altmetric MCPs can help.

The post Show your sources: building verifiable, citable AI agents with MCP appeared first on Digital Science.



from Digital Science https://ift.tt/TXdoacg

No Gold Standard: Measuring Success in Medical Affairs

To understand true publication impact and influence patient outcomes, Medical Affairs teams must have their own benchmarks

Compass Points: The Future of Medical Affairs is a series exploring the strategic challenges facing Medical Affairs teams in today’s communication landscape—and the tools that will help them get it right.

The best publication strategy is like a treatment plan: bespoke

In the past, journal citations served as the primary metric for measuring publication impact. Citations all but guaranteed a share of voice and influence amongst key opinion leaders and healthcare practitioners; they were also a simple, clear metric to share upward, proving research impact and justifying the allocation of resources.

Today, journals have come to occupy a different place in the Medical Affairs community. They still confer legitimacy, but they’re not the only way to have an impact; they aren’t even necessarily the most appropriate channel through which teams can or should disseminate information. 

How scientific information travels can be measured in both scientific impact and real world impact. Scientific impact comprises long-tail, more static measures such as citations and subsequent policy changes tracked over the course of months or years. But real-world impact—how a publication influences thought, conversation and even behavior—can be observed in how information ripples through other more immediate channels, like social or broadcast media and forums. 

Capturing an accurate picture of how a publication has performed requires a view of both. This holistic view allows teams to accurately benchmark performance, measure impact, and demonstrate value to stakeholders, supporting the overarching goal of improving patient outcomes.

The role of the journal has changed

While journals still heavily inform the provision of healthcare alongside clinical guidelines and regulatory bodies, they are not the only place members of the life sciences community can encounter and learn about new research.

As scientific information has come to travel on more horizontal, peer-to-peer channels, such as social media or podcasts hosted by trusted key opinion leaders, practitioners are able to learn about and interact with new research outside of the journal publication and conference cycle. This makes it easier for HCPs to stay on top of relevant research, and to quickly sift through the studies that are relevant to their clinical practice. This is why, depending on the therapeutic area and the goals of a given publication launch, Medical Affairs teams may find they gain more traction by diversifying to non-traditional channels. 

But in order for publication planners to take advantage of this reality—to optimize distribution across geographies, channels and a variety of timescales—requires dynamic, granular data that is consistently tracked through time. And as journals have come to form only part of the life sciences research diet, Medical Affairs teams have been left without a single, strategic reference point both for forward-planning and post-publication performance reporting. 

As a result, teams find themselves in a familiar position: unable to reliably demonstrate impact, defend decisions to stakeholders, or to quickly iterate for later distribution plans. This can undermine a team’s efficacy, and ultimately delay or limit influence on patient outcomes. 

What teams lose without a consistent benchmark

Benchmarking plays a critical role in publication planning. It allows teams to reference both the performance of earlier publications and the work of competitors, and to learn, in real time, what is working and what is not. Without benchmarks, it is impossible to know what is reasonable for research to achieve, and therefore, impossible to contextualize impact and prove a return on education. 

But benchmarking is also one of the most laborious parts of the publication planning cycle. The process of consistently benchmarking, tracking, reconciling and cleaning point-in-time data—from social media, journals, podcasts, conferences, magazine articles, and more—can take teams weeks of work. And because of the fragmentation inherent to the process, all of this work, ultimately, still may not be able to capture the nuance of a publication’s impact.

The knock-on effect is that teams are unable to design strategies which are optimized for a given therapeutic area and to meet specific performance goals, such as social media engagement. This undermines the team’s ability to demonstrate that strategic objectives are met and can limit the diffusion of information into communities that could benefit from it. 

This cycle repeats; without live, ongoing benchmarking, it’s impossible to see what’s changing in the competitive landscape and react to it. 

In recent years, the Medical Affairs community has matured dramatically with regards to its use of data. Teams rely heavily on analytics and data-driven decision-making. They know what data is available to them and how they can use it to inform publication strategies. 

But the tools available for assembling and parsing this data haven’t kept pace. Even as many teams embrace the use of general-purpose AI, the output is often unstandardized and non-reproducible—what AI surfaces today may be different from what it surfaces tomorrow, so teams can’t be sure they’re comparing like with like. In other words, teams gain speed, but not certainty.

This is what Compass by Dimensions was created to address. It brings together both traditional and alternative metrics so teams can easily benchmark against internal and competitor data, track publication performance through time and across channels in a standard and simple way, allowing to better demonstrate value, influence therapeutic behavior, meet education objectives, and trace real-world research impact. As a result, teams can move from publication strategies which are fundamentally reactive to those which are proactive, and as a result, better able to meet strategic objectives.

Figure 1: Workspace performance. Total number of attention events tracked across supported sources.

Tapping into the discussions that matter

Improving patient outcomes is the result of a confluence of events: research must be carried out, written up, disseminated, and then found by the relevant policy makers or healthcare practitioners to stand a chance of driving real-world impact. That means Medical Affairs teams need to look at both formal and peer-to-peer channels to measure influence. 

Compass is driven by data sources from two leading services in the scientific and research community, Dimensions and Altmetric. 

Dimensions hosts one of the largest collections of interconnected global research data, re-imagining research discovery with access to grants, publications, clinical trials, patents and policy documents all in one place. This data source provides a robust view of traditional channels.

Altmetric is a leading provider of alternative research metrics, helping everyone involved in research gauge the impact of their work. Altmetric searches thousands of online sources including social media, revealing where research is being shared and discussed—and the sentiment of that discussion. This is where real-world impact manifests first. 

In bringing these two data sources together, Compass allows teams to track the whole publication attention lifecycle, from social media posts minutes and hours after publication, all the way to citations and guideline mentions years after it was published, all reliably benchmarked through time. 

Teams need only set up a benchmarking dashboard to define which internal and competitor publications they want to track. Then the dashboard can be referenced any time for a simple view of a publication’s performance. Instead of the heavily manual work that teams had to endure before, Compass brings precise and consistent performance measurement and real-world benchmarking across assets or disease areas, giving teams the evidence needed to understand what’s working and where science is influencing practice.

Figure 2: Total mentions across domains.

Democratizing data-driven publication strategy

Medical Affairs teams are too-often forced to rely on costly and slow agency relationships to understand how their publications are performing. Teams who can take these tasks in-house with tools such as Compass will amplify the efficiency and agility with which they can work. 

Compass performs three functions which are critical to creating a truly bespoke, data-driven publication strategy

  • Unified, easy-to-reference metrics: Compass allows teams to quickly and consistently track performance through time to understand if a publication has reached the right people.
  • Custom benchmarking: Teams can benchmark publications against their own portfolio’s historical performance, establishing what a realistic result looks like. This exercise can also encompass competitors and specific therapeutic areas, helping teams identify opportunities and learnings. The task of surveying channels and creating internal and external benchmarks would have taken weeks before—now, once a dashboard is created, it takes just a few minutes to check.
  • Self-service interface: Compass creates shareable, stakeholder-ready visuals to clearly demonstrate the direct impact of work to decision-makers and support resource discussions with concrete data points. Teams can cut and re-cut data in a few minutes, saving the time, cost and hassle of looping in an agency every time a new report is needed.

Instead of waiting for an agency to return a report, or spending hours parsing data only for it to immediately stale, Compass places the power of real-time insights in the hands of the people best placed to wield it, streamlining the distribution process and supporting decision-makers with clear targets and performance measurement. 

Use case: Moving from lagging indicators to real-time feedback

Take a mid-size oncology-focused biopharma preparing to launch a publication for a second-line therapy. The Medical Affairs team’s usual process might take three to four weeks per reporting cycle: they would need to pull citation counts from one system, social and news mentions from another, then manually reconcile both in spreadsheets. If socials spiked after review had concluded on a given platform, that data would be missed. This means by the time a report reached leadership, the data would already be stale, and there would be no consistent way to benchmark the publication against prior launches in the same therapeutic area, because citations move too slowly and social media moves too quickly. 

For this team, Compass is designed to remove the manual burden of the benchmarking and tracking process. The team would first set up their benchmarking dashboard, tracking both the performance of previous portfolio publications and those of competitors working in specific therapeutic areas. Then, it would take just a few minutes to monitor performance: the team would easily be able to track journal citation activity against historical norms for the launch stage alongside engagement on clinician-focused social platforms and podcasts. 

Teams would be able to see where the research was finding the most resonance and quickly take action on that basis, for example, redirecting a portion of dissemination budget towards channels demonstrating traction. The consistency of this data also makes it easy to keep leadership informed; teams can slice and share reports from their Compass dashboard to illustrate what’s working and what isn’t. 

For teams that have previously relied on a single lagging metric, having a live, multi-channel benchmark can turn publication planning from a retrospective, best-guess exercise into a real-time, data-driven strategy.

The future of Medical Affairs is here—it’s time your strategy caught up

The proliferation of scientific information through popular channels is a good thing. 

That new scientific and medical information can reach much wider audiences through a variety of channels undoubtedly has a net positive impact on patient outcomes. Patient and rare disease advocates, healthcare practitioners working in remote or underfunded areas, and even patients themselves all benefit from having access to the cutting edge science that will shape the future of medical care. 

But the result of this proliferation has been a significant challenge for publication planners—no one could’ve predicted the rise of peer-reviewed podcasts. Now, teams need to take all of these data points into account when measuring performance and planning future publication launches. 

Manually compiling point-in-time data is a time-intensive process that puts publications at a disadvantage and undermines strategic and patient outcome objectives.

The right strategy is one entirely specific to a given publication—it must be tailored to relevant journals, audiences and channels, and these may all change through time. To date, with this level of nuance, it wasn’t possible for teams to keep up—not at the level of granular detail that could shape truly powerful publication strategies. 

Compass turns that complexity into an opportunity. 

The post No Gold Standard: Measuring Success in Medical Affairs appeared first on Digital Science.



from Digital Science https://ift.tt/2B3QOuq

From Writing the Rules to Building the Tools: Responsible Research Assessment in Practice

What does responsible research assessment actually ask of the people who build the infrastructure, rather than those who write the policy? In this post, Steven Hill traces the arc from DORA and the Leiden Manifesto through to the Barcelona Declaration, and sets out what the principles mean in practice for the tools that describe, discover, and measure research.

More than a decade ago, one of my first tasks in a new job was to advise on whether the organisation I had just joined, the Higher Education Funding Council for England, should sign the San Francisco Declaration on Research Assessment (DORA). We did, as a founding signatory, and it was the right decision. DORA’s central claim is that metrics, especially the journal impact factor, should not stand in as a proxy for the quality of an individual piece of research. As well as being right, that principle was an important signal that the UK’s national research assessment process was not taking a reductive approach to research quality.

What has stayed with me from that period is not the signing, but what came after. Alongside committing to an expanding set of principles building on DORA, the research system needs the patient work of turning those principles into reality. The Metric Tide, and its follow up seven years later, were in large part an attempt to take that problem seriously, and coined the term ‘responsible research assessment’, which labels the movement. The arc from DORA through the Leiden Manifesto, the Metric Tide, the Hong Kong Principles, the Coalition for Advancing Research Assessment (CoARA), and, most recently guidance from the Global Research Council (GRC), is the story of the global research community moving from declaration to implementation. The GRC, which brings together the heads of science funders from around the world, has provided funders with both tools to assess their own performance and a practical guide to making the changes needed in their practice. And the SCOPE framework for research evaluation offers a process for thinking through responsible research assessment in any evaluation context.

I find myself thinking about all of this again, but from an unfamiliar direction. For most of my career I have been a policy-maker, writing the principles and fretting about whether anyone is following them. At Digital Science I now look from a different direction: the building of the tools through which research gets described, discovered, and measured. Both setting the policy environment and helping to shape the tools bring power and responsibility, but the potential and the pitfalls are different.  The shift in vantage point raises a question: what should responsible research assessment ask of the people who build the infrastructure?

What the Principles Ask For – and What They Don’t

It is worth being clear about what responsible research assessment is, because it is easily caricatured. It is not a rejection of measurement, and it is not a plea to return to pure peer review uninformed by data. Read across DORA, the Leiden Manifesto, the Metric Tide, the Hong Kong Principles, and the CoARA agreement, and a consistent core emerges. Assessment should rest primarily on qualitative, expert judgement, with peer review at its heart, supported, not supplanted, by the responsible use of quantitative indicators. It should judge the work rather than the venue it appeared in. It should recognise the genuine diversity of what researchers produce and do: not only papers, but data, software, mentoring, peer review, public engagement, the often invisible labour of the people who make research possible. And it should be sensitive to context, to discipline, to career stage, and honest about its own limitations. Finally, as emphasised by the SCOPE framework, it is also important to critically reflect on whether evaluation is needed at all.

It is also fair to say that commercial entities in the research evaluation space are often criticised in discussions about responsible research assessment. The Leiden Manifesto asks that the data and the methods behind indicators be kept open and transparent, so that those being evaluated can verify them. CoARA goes further, calling for the research community to retain ownership and control of the infrastructure and the criteria used to assess it, and is openly wary of proprietary “black boxes”. The most recent Metric Tide review is blunt about the harm that commercial university rankings—built outside the academic community—continue to do to research culture.

Some of these critiques can be valid, although there are real practical challenges in realising total community ownership of data and infrastructure. And comparative analytics, well constructed and appropriately used, have a place in benchmarking universities. There is also the question of how commercial providers respond to responsible research assessment. The tools and the data are not going away; the question is whether they pull in the direction of the principles or against them. That is the real issue, and it should be the focus of the people who build the infrastructure, whether commercial or not, alongside the people who write the policies.

Why Openness Comes First

At Digital Science, colleagues here have been wrestling with this in public, through the lens of the Barcelona Declaration on Research Information. The Declaration’s first commitment is to make openness the default for the research information we use and produce—the records of who did what, where the money went, how outputs and contributions connect to one another—and to support the shared, open infrastructures that hold it. Writing on this blog, our CEO Daniel Hook has made the case that researchers have a fundamental right to access the metadata about research, and that the data used to evaluate academics should be transparently available and reproducible. He also argues that there are questions of assessment and measurement that will need data that is costly or complex to collect, and that openness might not be possible in this case. I think considering the balance and tension is the right direction, and it is worth dwelling on why, because open research information is the hinge on which the whole argument turns.

The responsible-metrics principles are simply not achievable on top of closed, unverifiable information. You cannot ask people to trust an assessment built on data they are not allowed to see. Open research information is the precondition, not an optional extra. But openness on its own is not enough. My colleague Simon Porter has written, again on this blog, about our responsibilities as consumers of metadata, not just producers of it. Use of research information needs to take into account the context in which it was generated, its provenance, and the extent to which the sources of information can be trusted, not just its availability. Information that is not accurate or appropriately contextualised can disrupt human judgement rather than support it. Simon also rightly notes potential equity concerns, where the metadata rich get privileged over the metadata poor, undermining the diversity and inclusion principle inherent in responsible assessment. He also notes that, as well as the responsible use of research information, responsible collection of data is also important.

Putting Principles into Practice

How does a commercial research infrastructure provider understand its role in supporting responsible research assessment? Rather than consider Digital Science’s products one by one, I want to focus on the principles of responsible research assessment and highlight examples where our tools and other options are aligned.

Broadening what counts. Research is more than journal articles, and the infrastructure has to be able to see and recognise a broader range of outputs. Being able to give a dataset a persistent identifier and a home, to surface software and preprints and policy documents alongside papers, to connect grants and patents and clinical trials into a fuller picture of a contribution is at the heart of responsible assessment. Digital Science tools such as Symplectic Elements, Figshare, and Dimensions, and the tools and work flows that they enable, are useful here precisely to the extent that they make the diverse outputs visible and creditable.

Supporting judgement rather than replacing it. The most valuable thing a system can do is not to produce a number, but to assemble as broad a view of the available evidence, so that human beings can exercise judgement well, and a researcher can tell their own story. Dimensions includes a range of tools that enable decision-makers to access clear summaries of the data and evidence that they need. Research information systems, such as Elements, that support narrative and evidence-based CVs, and that spare people the indignity of re-keying the same information into yet another form, are doing something genuinely in the spirit of the reform. The recently introduced CV import capability in Elements contributes directly to this objective.

Many dimensions, not one. When Altmetric first appeared, its real purpose was not to provide a new “score” but to emphasise evidence of broader contributions beyond those measured through citations. Evidence of attention in policy documents, in the press, in clinical guidance tells you something a citation count cannot. Links between publications and patents and policy documents in Dimensions also provide this richer picture of research. Outside of the Digital Science product line, Overton also provides data on the rich connections between research and policy.

Transparency and context. This is where the Barcelona Declaration is important, and Digital Science’s Open Principles set out how we work to align our tools with its aims. Making core elements of the Dimensions and Altmetric datasets freely available sits at the heart of these principles, alongside our commitments to work with the research community, and to openly publish our thinking and research. For example, where Dimensions data are used for assessment purposes researchers and their employers can check and verify the data. Our data sits alongside other open sources such as Crossref and DataCite and persistent identifiers like ORCID and ROR, key parts of the open responsible research assessment infrastructure. OpenAlex also offers fully open information as a secondary aggregator, overlapping in some areas with Dimensions.

I have spent enough time on the policy side to be wary of believing that any of this can be solved by better tools alone. Responsible research assessment is about behaviours, norms and incentives as much as it is about systems and infrastructure. And the choice isn’t between commercial infrastructure and community-owned systems. What matters is that infrastructure is built and used in a way that supports human judgement, broadens what we value, and submits itself to transparency and scrutiny. This is what responsible research assessment asks of those who build the infrastructure, and should inform everything we do at Digital Science.

The post From Writing the Rules to Building the Tools: Responsible Research Assessment in Practice appeared first on Digital Science.



from Digital Science https://ift.tt/9kBIOJf

Why Your AI Agents Are Only as Good as the Knowledge Behind Them

The race to deploy AI agents is accelerating, but most organizations are still building on sand. A new Gartner report suggests that the key to building reliable AI agents is a “context layer”.

According to Gartner’s latest research, 42% of enterprises plan to deploy AI agents by the end of 2026, and AI agent spending is expected to grow from 22% to 31% of total AI budgets in just one year1. Despite this wave of investment, only one in five organizations report that their GenAI tools are delivering significant value. Hallucinations, limited impact, and unpredictable behavior remain stubbornly common.

The problem, Gartner argues, isn’t the models, but rather what surrounds them.

The Missing Layer

Behind every reliable AI agent is something Gartner now calls a “context layer”— a dedicated architectural component that curates, organizes, and delivers the knowledge an agent needs to act intelligently. Without it, agents are left processing noisy, poorly prioritized data, making expensive errors and producing outputs that can’t be trusted or traced.

Gartner is unambiguous about the stakes: by 2027, organizations that prioritize semantics in AI-ready data could increase their agentic AI accuracy by up to 80% and reduce costs by up to 60%. The context layer is no longer an optional refinement — it is the necessary foundation.

And yet this layer cannot simply be purchased. No vendor offers it out of the box. It must be engineered, assembled from services, capabilities, and custom modeling that together transform an organization’s tacit knowledge into something AI agents can actually use.

Three Components, One Foundation

As stated in the report, there are three interlocking components that make up this ‘context’ layer: semantics, operational state, and provenance. Together, they form a pipeline that allows agents to retrieve the right information, organize it coherently, and act on it with accountability.

Semantics: Meaning, Not Just Data

Semantics is the component most organizations are missing, despite it being the one with the greatest leverage. Gartner finds that organizations implementing semantic modelling such as ontologies and knowledge graphs, are 2.2 times more likely to achieve high effectiveness in AI data engineering, however, only 40% of organizations have done so. 

Semantics means representing your organization’s knowledge—business entities, rules, policies, relationships, metrics—in machine-readable form. This allows AI agents to interpret what something means in context and execute an action based on that context, not just pattern-match on keywords. Without this layer, even the most sophisticated agent is, in effect, guessing.

This is precisely the domain where metaphactory brings long-standing proven capability. metaphactory by metaphacts, a Digital Science solution, is a knowledge graph platform enabling organizations to build and maintain rich semantic models for over a decade—connecting business glossaries, ontologies, and data products in ways that AI agents can directly leverage. For organizations serious about agentic AI, a robust semantic foundation isn’t a future aspiration; it is a prerequisite.

Operational State: The Right Information at the Right Time

While semantics provides meaning, your operational state provides situational awareness. AI agents need access to current, accurate information about the entities and processes they’re acting on beyond just snapshots, such as up-to-date information on customers, datasets, experiments, publications and suppliers. 

For research-intensive organizations, this is particularly acute. The ‘operational state’ of a research environment spans live datasets, ongoing experiments, researcher expertise, institutional repositories, and the evolving landscape of published science. Digital Science’s portfolio—including Dimensions, Altmetric, and Figshare—represents exactly this kind of curated, continuously updated operational knowledge. Rather than building this knowledge from scratch, organizations working in research and innovation already have access to a pre-assembled foundation.

Gartner also highlights the Model Context Protocol (MCP) as the emerging standard for connecting agents to operational state efficiently and securely. Dimensions, Altmetric, and metaphactory already support MCP, reflecting a broader conviction that research infrastructure should be designed to meet agents where they are, not retrofitted after the fact. As adoption of the protocol grows across the industry, having well-structured knowledge accessible through it will matter more, not less.

Provenance: Trust Through Traceability

The third component—provenance—is what makes agentic AI governable. It encompasses the systematic tracking of data lineage, agent decisions, actions, outcomes, and feedback across the full lifecycle of AI operations.

For research organizations, publishers, and funders, provenance isn’t merely a governance checkbox. It is central to the integrity of the work itself. Reproducibility, accountability, and the ability to audit AI-assisted conclusions are not simply peripheral concerns; they are defining ones. Gartner notes that 74% of organizations recognize that data governance tools are essential to operationalizing AI governance, yet robust provenance mechanisms remain rare in practice.

Digital Science’s longstanding commitment to open, traceable research infrastructure, including persistent identifiers, transparent data lineage and open metadata, gives research organizations a natural head start on this component. The challenge is connecting these capabilities explicitly into the agentic architecture, so that every AI-assisted decision can be traced back to its sources and reviewed.

Research Intelligence as a Context Layer

There is a broader framing worth making explicit here: for organizations operating in research, science, and innovation, the context layer is not merely a technical architecture problem. It is, at its core, a research intelligence problem.

The tacit knowledge Gartner describes—the organizational understanding that must be made machine-readable for AI agents to function—is, in a research context, the accumulated intelligence of a scientific community: what has been discovered, by whom, with what methods, validated how, and applied where.

We have spent over a decade building infrastructure that captures precisely this kind of knowledge at scale. The shift to agentic AI doesn’t make that infrastructure less relevant—it makes it more so. The question is no longer just “can researchers find the right information?” but “can AI agents, acting on researchers’ behalf, find, interpret, and act on that information reliably and accountably?”

The answer depends entirely on the quality of the context layer underneath.

What This Means in Practice

For R&D leaders and data and analytics leaders, the practical implication is this: before asking which AI agent to deploy, ask what context layer you have in place to support it. Gartner’s advice is to start with high-value use cases rather than attempting a comprehensive build all at once—iterate, demonstrate outcomes, and expand. That is sound counsel. But iteration without a semantic foundation, without right-time data access, and without provenance mechanisms will simply produce faster failures.

The organizations that will lead in agentic AI are not those that move fastest to deploy agents. It is the organizations that invest earliest in the knowledge infrastructure that make agents worth deploying.

Digital Science is working with research organizations and data-intensive enterprises to build the context layers their AI strategies require.


  1. Gartner. (2026). The 3 core components of the context layer for AI agents. [Research Note/Report]. https://www.gartner.com/document/ [G00848874]

The post Why Your AI Agents Are Only as Good as the Knowledge Behind Them appeared first on Digital Science.



from Digital Science https://ift.tt/1isBFqy

Featured Post

Grounding AI Agents: literature vs. structured databases in the biopharma data stack

Sharing perspectives from the hubXchange 2026 Roundtable: “Assembling the data stack: what pharma needs from external knowledge in the age ...

Popular