The next era of AI in drug discovery

As AI moves from summarizing papers to generating scientific claims, biopharma faces a hard question: how do you trust an answer with no way to trace it back to the truth? Digital Science’s Mark Hahnel unpacks the shift at hubXchange 2026.

Keynotes & insights from AI in Drug Discovery hubXchange 2026

On September 9, 2026, leaders across the biopharmaceutical and artificial intelligence sectors gathered in San Francisco for the AI in Drug Discovery hubXchange. Digital Science’s Mark Hahnel delivered a keynote address examining how AI is altering scientific inquiry and biopharma R&D. Moving beyond basic task automation, Hahnel articulated a future centered on data provenance, agentic workflows and the assembly of an industry-wide data substrate.

Read on for a recap and key takeaways from Hahnel’s keynote.

Moving beyond ‘AI as a Tool’ to agentic workflows

Scientific research has entered the ‘4th paradigm’, where massive volumes of data must be readily available and actionable. However, the current velocity of AI adoption is accelerating at a rate that introduces friction into traditional organizational workflows. Rapid theoretical milestones highlight this momentum, such as OpenAI publishing proofs for the complex Navier-Stokes fluid mechanics equations on Twitter. The tension between independent mathematicians leveraging Codex to tackle similar challenges, and claims of unethical scooping, demonstrates the need for tools that protect IP while supporting research advancement.

Large corporations are embracing AI, transitioning from basic productivity aids to deploying AI for generating new scientific knowledge and discovering novel drugs. Room consensus at hubXchange aligned with statistics indicating over 50% enterprise adoption across major organizations. The biopharma industry has officially progressed past using isolated AI tools and entered the next phase: managing siloed data alongside specialized domain models and autonomous agentic workflows.

Offline & local models: safeguarding intellectual property

As enterprise adoption deepens, maintaining strict data security and protecting early-stage IP remain paramount. To prevent proprietary research from being inadvertently exposed or ‘scooped’ through public cloud chat windows, biopharma companies are prioritizing local models and offline data architectures.

Dedicated tools like Digital Science’s Papers AI address this demand by keeping enterprise data, local models, and analytical routines fully offline, ensuring researchers can leverage modern AI capabilities without compromising security.

The Provenance Crisis & Data Trust

As generative systems output claims at scale, biopharmaceutical organizations face a foundational challenge: Where did this data originate, and how was this specific claim validated? Grounding claims in verifiable truth is critical because core human facts reside outside the latent weights of Large Language Models (LLMs).

To achieve full traceability, organizations must back up every statement and derived insight. Hahnel highlighted the framework detailed in Digital Science’s FAIR data playbook for Pharma white paper as an essential roadmap for establishing structured, trustworthy data environments. Regulatory compliance necessitates adhering to the FDA + EMA Guiding Principles of Good AI Practice in Drug Development, which require tracking the explicit source, raw underlying data, and precise timestamps for all AI-assisted findings.

Constructing & deconstructing papers for machines

Building robust drug discovery models requires looking beyond high-level literature summaries. While cheap and accessible methods exist—such as using models like Claude to ingest titles and abstracts from PubMed—true drug discovery demands deep full-text extraction. Full text is essential to extract vital scientific nuances, including detailed methods, experimental edge cases, figures, and direct scientific contradictions.

Navigating this domain requires working within a fragmented publisher landscape, where the top 5% of publishers account for approximately 61% of all scientific publications. Existing pharma licensing agreements provide a pathway to deconstruct and reconstruct scientific papers into machine-ready structures, enhancing internal proprietary data.

Data must be structured once across workflows so it can be continuously reused rather than repeatedly extracted. Maintaining these comprehensive global databases requires continuous operational maintenance; for example, maintaining Dimensions‘ global patent database requires a workforce actively liaising with patent offices worldwide to correct inaccuracies and guarantee precision.

Four core techniques for structured data extraction

To derive locally verifiable statements and establish end-to-end data provenance, four primary computational techniques are being actively deployed:

  • Mapping: Leveraging LLM-driven ontology mapping to harmonize disparate scientific terminologies across domains.
  • Graphs: Building dynamic knowledge graphs that represent evolving biological relationships and entities.
  • Triage: Implementing just-in-time triage and filtering to parse incoming streams of scientific literature efficiently.
  • Extractors: Deploying agentic extractors designed to pull out claims, experimental methods, biological entities, and explicit relationships directly from full text.

These techniques allow organizations to extract claims and harmonize them so they are composable with internal proprietary extensions, Electronic Lab Notebooks (ELN), and existing R&D workflows.

Conclusion: the substrate is the work

The primary takeaway from Mark Hahnel’s presentation is clear: while foundational models and agent frameworks will continuously improve, the ultimate value lies in the data substrate beneath them. Grounded scholarly inference requires generating answers built upon verified external data seamlessly combined with internal enterprise assets.

Building this substrate requires deep collaboration across biopharma, biotech, academic publishers, and AI tooling providers. Models and agent frameworks will continue to evolve, but establishing the underlying, composable data substrate is the foundational work that the entire field must build together.

Ready to build the data substrate your AI strategy depends on? Explore how Digital Science’s enterprise solutions help biopharma organizations turn siloed data into trusted, structured, AI-ready assets.

The post The next era of AI in drug discovery appeared first on Digital Science.



from Digital Science https://ift.tt/Zg0Jmiy

In Life Sciences, data integrity is non-negotiable

In the age of AI, research intelligence has to begin with trusted data. The Inside Our Data series explores the foundational data infrastructure that makes trust possible. 

The cost of an answer no one can explain

Enterprises rely on research intelligence to set and meet strategic objectives; intelligence is what makes an enterprise competitive. Research intelligence, as a concept, isn’t new. What is new is the enterprise’s necessary reliance on huge volumes of data—and on AI-driven insights and analysis which take that data as truth. 

Enterprises are moving fast to embed AI into research, analytics, and strategic planning. Far from the zeitgeisty pilots which were largely based on frontier model usage, AI-driven intelligence now comprises foundational infrastructure that informs how decisions are made. 

But for the advances and benefits this technology has already brought about, it has also brought risk. Research and AI-driven intelligence are only as trustworthy and defensible as the data they use, and not all data can withstand the necessary scrutiny of independent review.

This can pose an existential threat to research enterprises operating in regulated industries such as Life Sciences. As enterprises continue to evolve and embrace powerful new technologies, it’s more important than ever that their underlying data can stand up to audit. 

In this article, we’ll look at what happens when enterprises lack data integrity, what “trustworthy” data actually means in Life Sciences, and how the right infrastructure can fortify enterprise data in the world of AI.

What happens when the data underneath a scientific conclusion can’t be checked?

In 2020, two COVID-19 studies were published—one in the Lancet and one in the New England Journal of Medicine—using data from Surgisphere, a little-known analytics firm. 

But soon, there was a problem: Surgisphere refused to release its data for an independent audit. Both studies were retracted nine days apart.

The Lancet study claimed hydroxychloroquine increased mortality risk in COVID-19 patients—a finding that prompted the WHO to briefly pause a hydroxychloroquine arm of its global Solidarity trial before the retraction. The NEJM study was also retracted, but kept being cited long after: a Journal of the American Medical Association Internal Medicine analysis found 652 verified citations, with more than half of them occurring at least three months after the retractions took place. 

This incident illustrates what is at risk when data can’t be audited. We don’t know why Surgisphere wouldn’t release the data. Maybe it was all fabricated, maybe it wasn’t. There’s no way to know. But it doesn’t really matter. Data that cannot be audited is contagious. Bad data doesn’t stay where it started. It moves into papers, then models. Bad data has always been contagious. AI gives it a much higher reproduction rate.

It’s likely that the initial retractions were costly and frustrating for the firms who carried out the studies. But this incident also contributed to a wave of inaccuracy in critical research areas. Data that can’t be audited can halt clinical trials, knock percentage points off a stock valuation, cause lasting reputational damage, and most seriously, negatively impact the lives of real people. 

What “trusted data” really means

When it comes to defining what makes good data, enterprises aren’t starting from square one—Life Sciences has already formalized what “trustworthy” data means.

Regulators have relied on ALCOA—Attributable, Legible, Contemporaneous, Original, and Accurate—since the 1990s to assess data integrity in clinical and manufacturing contexts. More recently, guidance from bodies like the Medicines and Healthcare products Regulatory Agency and the World Health Organization extended this into ALCOA+, adding four further requirements: data should also be Complete, Consistent, Enduring, and Available. It’s a checklist built around whether a record can be verified after the fact, and it’s still required for trustworthy data today.

The related framework, FAIR—Findable, Accessible, Interoperable, and Reusable—addresses a different but equally consequential point of failure. Introduced in 2016, FAIR has had a substantial impact on how research-generating organizations think about their data: not just whether it exists, but whether it can be found, retrieved under clear terms, and reused with confidence in its provenance. FAIR doesn’t require data to be open to all—a dataset behind a paywall or access agreement can still be fully FAIR-compliant. However, in order to satisfy the requirements of the framework, it needs to be made available, in a FAIR manner, to the people who would make assertions on that data. 

The Surgisphere retractions occurred because the underlying datasets were non-compliant with these frameworks. The data wasn’t available or accessible for audit, which meant we also couldn’t know if it exemplified the necessary integrity that made it suitable for use in research.

These frameworks comprise a non-negotiable baseline for Life Sciences research enterprises, but there remain grey areas which can have unintended effects on the quality of research datasets. For example, an open dataset which is seemingly FAIR and ALCOA+-compliant could be skewed toward whichever countries or funders proactively volunteer their data. In a regulated environment, this isn’t enough; passing an ALCOA+ or FAIR checklist doesn’t tell you whether a dataset is representative—curation, applied on top of these frameworks, can correct for that skew. 

Data infrastructure designed to accommodate investigation

Data curation refers to the ongoing process of ensuring that data is complete and representative—a process that requires human judgment and relationships to execute. This is a foundational tenet of the datasets which comprise Dimensions by Digital Science, one of the world’s largest research and funding data repositories.

Dimensions was built around the idea that research intelligence is only useful if it can be traced across the full lifecycle it describes—not only publications, but the funding, trials, patents, and policy activity that surround them. Dimensions datasets span six linked content types: more than 165 million publications, 8.2 million grants, 74 million research datasets, 180 million patents, 976,000 clinical trials, and 2.5 million policy documents, all cross-referenced with the others. 

A model surfaces a promising area of research. Don’t just take the answer. Ask:

  • Who funded it?
  • Which researchers produced it?
  • What publications followed?
  • What datasets underpin them?
  • Were patents filed?
  • Did it progress into clinical trials?
  • Did it influence policy?

This structure enables attributability and originality under ALCOA+: publication records are enriched through full-text indexing and linked back to direct publisher partnerships, Crossref, PubMed, and other authoritative sources. The grant data comes from more than 700 funders worldwide, sourced by data experts directly from funder organizations wherever possible. Clinical trial records are pulled directly from official registries spanning every major region, so status, sponsors, and outcomes reflect the authoritative record rather than a secondhand summary. Patent data is provided by IFI Claims, curated and normalized by Digital Science teams.

Every record carries a persistent identifier and a link back to its original source, so a grant, publication, or patent is Findable and its provenance is never in question. Records are Accessible under clear, documented terms—whether that’s open data or a governed connection through a licensed platform, so users always know what they’re looking at and where it came from. And the cross-referencing between content types is what makes the data Interoperable and Reusable in practice: a grant can be traced through to the publications it funded, the datasets and patents those publications generated, and the clinical trials or policy documents that followed. Research across more than 100 countries and every major discipline reduces the blind spots that come from a literature-only view or a single-region dataset. This is data that can be audited—and that enterprises can trust to drive the decisions they make. 

The final step to unlocking truly powerful and trustworthy intelligence is ensuring this data infrastructure is in sync with enterprise-specific ontologies. Pairing trusted data with semantic definitions lays the groundwork for life sciences enterprises to more safely rely on AI-driven research intelligence in the years to come.  

Building trusted foundations for future AI implementations

The Surgisphere debacle exemplifies the failures that AI-assisted workflows now risk automating at scale: fluent, confident outputs based on data that doesn’t meet industry standards. Today, AI-assisted workflows are increasingly embedded in how R&D and Medical Affairs teams triage literature, surface signals, and make decisions. The efficiencies and insights to be gained from this technology are unprecedented, but this also raises the stakes: an AI working from ungoverned data doesn’t just produce a bad answer, it can introduce existential risk. 

The best way to guard against such a failure—and set your company up for long-term success—is to take a two-pronged approach, pairing an enterprise-specific semantic layer, such as a knowledge graph, with data that is FAIR and ALCOA+-compliant. Digital Science offers knowledge graph infrastructure designed to grow with an enterprise via its proprietary technology, metaphacts

By defining a semantic layer, an enterprise sets the scope for the data that AI is able to access and defines the logical relations between defined entities. This means the model can only interpret the data it is given access to in the context of an approved series of rules. This mitigates the risk of hallucinations or logical failures, and makes it simple for auditors to interrogate the pathways that led to a certain output. This is how to ensure trusted data is treated predictably by trusted models.

The result is accurate intelligence with an in-built audit trail that enterprises can trust to stand up to independent audit. 

The bar for data integrity will keep rising

As AI becomes more embedded in R&D and Medical Affairs decision-making, so too will audits by regulators and internal stakeholders. Trusted intelligence starts with trusted data, and trusted data is best used in sync with foundational enterprise infrastructure.

Digital Science provides one of the world’s broadest collections of connected research intelligence—combining Dimensions, Altmetric, and IFI Claims to help enterprise organizations support analytics, strategic decision-making, innovation, and AI workflows that can be explained, audited, and defended. 

The post In Life Sciences, data integrity is non-negotiable appeared first on Digital Science.



from Digital Science https://ift.tt/QRGXa1F

A double-edged sword: the growing complexity of Medical Affairs publication performance data

The variety of channels and audiences that define scientific communications reach and engagement is growing. In turn, Medical Affairs teams face diversifying data sources and tools to assess publication performance. 

Compass Points: The Future of Medical Affairs is a series exploring the strategic challenges facing medical affairs teams in today’s communication landscape—and the tools that will help them get it right.

Even the most groundbreaking data cannot change clinical practice if never translated into action. As such, a fundamental purpose of scientific communications is to inform and educate on this new data, what it means, and how it can impact the real world. The challenge is how to do this effectively across multiple regions, channels, and audiences, and how to track success (or failure).

As the complexity of scientific communication scales, Medical Affairs teams rely on an expanding library of data sources and tools to analyze the performance of scientific communications tactics. Quantifying asset performance and impact directly informs strategy, and in turn, informs publication planning. We see that this feedback loop propagates the outcomes of tactical and strategic decision-making, whether these outcomes were desirable or undesirable.

Publication planning and performance feedback loop

The growing availability of data sources and tools used to define publication performance is a double-edged sword: capabilities increase, but so, too, does workload. Assessments performed in different settings, at different time points, with non-standardized queries may create inconsistency in those outputs contributing to strategic decisions about publications. The value of scientific communications can be efficiently captured by measurement tools, such as Compass by Dimensions, characterized by integrated sources, standardized data, and intuitive performance benchmarking.

Limitations become visible when publication performance data sources and reporting tools are siloed.

As part of Medical Affairs scientific communications planning and evaluation, asset performance directly informs publications strategy. A growing variety of data sources and analysis tools are now available. These help determine publications’ reach and engagement, and by extension, their impact.

Citation tracking tools hold continued relevance. In what may represent a highly manual process, pertinent altmetrics must first be defined, then followed over time. Social media listening offers publication performance insights from an altogether different channel. To assess proprietary (or competitor) abstracts, posters, and podium presentations, congress trackers of varying complexity are commercially available or developed in-house. Whether for conferences, publishers, or individual journals, both the type and availability of performance metrics vary widely. 

These examples are not comprehensive. As their variety suggests, publication performance data sources and reporting tools are often functionally siloed from one another. They must be evaluated in turn, and the readouts integrated, to generate a comprehensive snapshot. 

Being inherently decoupled, it follows that the data sources and reporting tools illustrated here will lack technical platform interoperability. Plainly stated, they don’t communicate. As such, they are limited in their ability to provide integrated readouts and a contextual story of scientific communications asset performance.

What does this mean for user experience and workload?

Across life sciences industries, the size, structure, and distribution of Medical Affairs and publications teams differ significantly. Scientific communications strategy may be defined within the Medical Affairs functional area alone, or within a cross-functional center of excellence or integrated evidence planning team.

Where data and reporting tools are managed by a group of colleagues, only by investing time and aligning their efforts can these contributors integrate findings into a cohesive performance narrative. If such coordinating and reporting activities are repeated on a monthly basis, for example, we begin to grasp the many people-hours required. In the present era of remote work, it’s likely that these team members do not work in the same physical space, or even the same time zone. Creating the impact story requires continuous touchpoints, further decreasing efficiency.

It is important to highlight this concept of the scientific communications impact story, as creating it is just one step in the process. Another key aspect is telling that impact story effectively to leadership and other key stakeholders. How are the publication performance data contextualized? What reporting content can decision-makers expect to see, and reliably? 

A holistic scientific communications performance overview, delivered on-schedule with consistent format, takes significant time and effort, whether the overview’s creator is a team or a single contributor.

In the case of a single contributor such as the publications manager or director, this colleague is solely responsible for the time-consuming, repetitive work of integrating increasingly complex data sources and tools. Expertise more impactfully invested in key project management and strategic activities is instead diverted to data analysis. The workload risks overwhelming that colleague.

Whether in (bio)pharma, biotech, or medtech organizations, this situation’s impact may be more acutely felt in publications teams serving multiple disease or product areas. In a further example, its impact is visible in small- and medium-sized life sciences companies, where publications colleagues may “wear other hats,” having broader role descriptions or functional responsibilities.

When publication performance insights are integrated from diverse sources, how does this influence their perception?

Building on this, publications teams are facing operational environments in which scientific communications performance assessment and reporting processes become overwhelming.

While these may be subject to formalized standard operating procedures, it’s more likely that practices fluctuate over time: team structures change, or publication types evolve. Inherent knowledge informs the work of integrating performance data from diverse sources, often depending on personal best practices. Processes become opaque, and as the risks of missing relevant data and of differing interpretations increase, reporting inconsistencies emerge.

Whether monitoring owned or competitor assets, publication performance reporting serves myriad purposes. These range from publication impact measurement, to downstream budget and strategy planning, to competitive intelligence. Performance reporting is meant to describe impact and value.

Should the integrated insights appear inconsistent, this perception affects stakeholders. It reflects negatively on the work and reputation of the publications or scientific communications team, the Medical Affairs team, or the integrated evidence planning team. Cross-functional partners or leadership may perceive the accumulated insights as unreliable, or even non-actionable. Over time, this hinders effective business decision-making, perceived department value, trust, and even individual working relationships.

Data integration workarounds that utilize generative artificial intelligence lack fidelity.

In the last three years, multimodal generative artificial intelligence (genAI) technologies have gained significant traction as data integrators. Their ability to instantaneously compare inputs, summarize findings, and create personalized outputs feels reassuring. With remarkable efficiency improvements, a single user can develop polished, on-brand content and dashboards in minutes.

GenAI technologies may represent a tempting solution to the challenge of publication performance data collected from such disparate sources and tools. This is especially true for life sciences organizations holding enterprise agreements that facilitate company-managed access to these technologies.

It is critical to balance the benefits of improved efficiency against the limitations of utilizing genAI as a process workaround to analyze and integrate publication performance data. Due to these technologies’ very design, they are neither able to consistently benchmark nor to track target performance metrics over time. As such, assessments remain snapshots that must be repeated according to stakeholders’ reporting requirements.

Hallucination and sycophantic responses are known challenges with the use of genAI. Outputs with publication performance data integration as their goal may be incomplete, factually incorrect, or biased. A genAI-grounded process still relies on the user to identify and supply trusted data sources. If pertinent metrics are missing, genAI-directed data integration processes cannot account for them. Alternatively, depending on how the user prompts the model, they may have the undesirable experience of hallucinated metrics or outputs.

The use of genAI to speed up integration of disparate, disconnected data sources should not come at the cost of insight fidelity. Rather, when artificial intelligence capabilities are paired with data analytics, reliable analyses require standardized, consistent data feeds from curated sources. When a publications team builds such analytics de novo, both the data sources and analytics outputs take time to verify and to trust.

Standardization and repeatability are key to successful publication performance assessment.

Capturing the value of scientific communications should not be held back by the repetitive work of reconciling disparate data sources. Nor should strategy-defining insights depend on workarounds, themselves subject to technical limitations. As well, it is worthwhile to consider the accumulated inefficiencies that these activities create for publications managers and teams.

Measurement tools that integrate data sources by their design unlock the power of user-defined search and tracking parameters. Meaningful insights are uncovered when these parameters are standardized and repeatable, tracking publication performance with consistency over time. When unique, Medical Affairs-relevant data sources come already embedded, it streamlines the work of uncovering scientific communications reach, engagement, and impact. This diversity of data is no longer an obstacle.

Compass by Dimensions captures these capabilities. Built on more than a decade of Dimensions and Altmetric data trusted by industry, it is designed to help overcome the challenge of data diversity. Compass combines publication and altmetrics into a single collaborative workflow, reducing inefficiencies, saving time, and simplifying how publications professionals and Medical Affairs teams benchmark, track, and manage publication impact and reach.  

Compass by Dimensions is developed by Digital Science, an AI-focused technology company that transforms fragmented data into unified knowledge assets, leveraging AI and Knowledge Graphs to deliver structured, actionable intelligence for high-value discovery and innovation. By combining unparalleled data depth and breadth with enterprise-ready AI technology, we help leaders confidently accelerate product life cycles and secure a decisive market lead.

The post A double-edged sword: the growing complexity of Medical Affairs publication performance data appeared first on Digital Science.



from Digital Science https://ift.tt/tuWEXl6

48th Edition of Global Scholar Awards | 26–27 September 2026 | Hong Kong, China - Cordis Hong Kong


Global Scholar Awards recognize remarkable individuals who demonstrate excellence in research, education, innovation, and leadership. The program promotes academic distinction, encourages groundbreaking discoveries, and honors contributions that improve communities, advance scientific understanding, and inspire continued progress across multiple disciplines worldwide.

Global Scholar Awards 🌟

Visit Our Website 🌐: globalscholarawards.com Nominate Now👍: https://globalscholarawards.com/doctor-awards-nobel-prize-scientists-award-nomination/?ecategory=Awards&rcategory=Awardee Contact us ✉️: info@globalscholarawards.com Get Connected Here: ================= Twitter : x.com/ScienceInventi1 Youtube : youtube.com/@nesinconferenceandawards4869 Pinterest : in.pinterest.com/scienceinventions/ Instagram : instagram.com/global_scholar_123 Linkedin : linkedin.com/in/global-scholar-awards-09664427b Blog : newscienceinventions2020.blogspot.com @WorldResearchAwards @GlobalScholarAwards #worldresearchawards #researchawards #researchexcellence #globalrecognition #academicawards #globalresearchawards #shorts #researchers #labtechnicians #awards #professors #teachers #lecturers #engineering

Dr. Wael Megid | Engineering | Best Researcher Award


 Global Scholar Awards 🌟

Visit Our Website 🌐: globalscholarawards.com Nominate Now👍: https://globalscholarawards.com/doctor-awards-nobel-prize-scientists-award-nomination/?ecategory=Awards&rcategory=Awardee Contact us ✉️: info@globalscholarawards.com Get Connected Here: ================= Twitter : x.com/ScienceInventi1 Youtube : youtube.com/@nesinconferenceandawards4869 Pinterest : in.pinterest.com/scienceinventions/ Instagram : instagram.com/global_scholar_123 Linkedin : linkedin.com/in/global-scholar-awards-09664427b Blog : newscienceinventions2020.blogspot.com @WorldResearchAwards @GlobalScholarAwards #worldresearchawards #researchawards #researchexcellence #globalrecognition #academicawards #globalresearchawards #shorts #researchers #labtechnicians #awards #professors #teachers #lecturers #engineering

Show your sources: building verifiable, citable AI agents with MCP

Model Context Protocol (MCP) is an open standard which connects LLMs with external systems. We discuss how new Dimensions and Altmetric MCPs ground LLMs in structured data, generating verified, citable results that research teams can trust.

The “Knowledge Gap” in AI

You can’t create cutting-edge research from stale, outdated information. And yet research teams are trying and failing to derive insights from standard Large Language Models (LLMs) trained on data that is months or even years old. A significant problem in research contexts where new information is constantly released, and hundreds – if not thousands – of new publications and reports are published daily. 

AI tools can take over time-consuming research tasks like competitor tracking, but connecting them to trusted data can take more custom coding and login/security setup than most teams have the resources for. When research teams lack the necessary technical abilities to hard-code, they are alienated from the successes of AI Research Integration. One in three R&D-focused enterprises say understanding and implementing AI tools is one of their biggest challenges.

With the Model Context Protocol (MCP), no team needs to miss out. In this article, we explain what an MCP is, why it’s governance-friendly, and explore how MCPs elevate strategy and innovation with Dimensions and Altmetric MCPs

What is Model Context Protocol (MCP)?

Large Language Models (LLMs) promise enormous potential. But the potential of these models has been stunted. This is because the data AI interacts with is siloed and trapped behind legacy systems. AI models are thus forced to use outdated training data to make their decisions, and can hallucinate when asked to reason over current scientific literature.

For example, a researcher exploring lung cancer treatments may ask an LLM to identify, based on all available scientific literature and past oncology trials, the toxicity risks of a promising drug. The LLM outputs an authoritative and neatly argued “green light” for the use of this drug, for this new application. What this researcher doesn’t realize is that not only has the LLM missed an influential dataset published last week, which evidenced harmful effects for patients with specific comorbidities, but also hallucinated a single decimal point in a critical dosage threshold. 

In a study by NVIDIA, 59% of respondents from pharmaceutical and biotech companies cited drug discovery and development among their top AI use cases.

Previously, if you wanted to integrate AI into an external service, this required lengthy custom implementations. Now, the Model Context Protocol (MCP) provides a standardized, plug-and-play protocol for AI applications to connect to external data sources in a structured, reliable, and permissioned way – similar to how a USB-C port allows external devices to connect to computers. This means autonomous agents (like Claude or ChatGPT) can query relevant data using natural language instead of complex code. 

With MCP, you can connect AI assistants like Claude, Cursor, VS Code Copilot and ChatGPT with external data sources and tools. This might include productivity tools (like Slack), development tools (like GitHub) and data and file systems (like Google Drive). 

Consider an engineering team trying to develop a new kind of lightweight battery for electric cars. Previously, when using LLMs to challenge and improve their prototypes, they would copy and paste their research into the chat window each session. By connecting AI to their data sources and tools via MCP, copying and pasting their research became unnecessary. Not only does the AI automatically query the relevant research mid-conversation, but it can also surface relevant context from a years-old study, buried in the organization’s Google Drive. This connection sparks the game-changing insight. 

MCP: the data flow your IT team will actually approve

Teams who primarily work via cloud services will be – perhaps acutely – familiar with the lengthy AI governance for research required for new SaaS tools. In these kinds of environments, AI governance requires the continuous real-time monitoring and risk assessment of every team member’s AI adoption. In a fast-paced research environment, this is a time- and resource-intensive undertaking. According to a 2025 AI-Ready Governance Report, organizations reported a 37% jump year-on-year in time spent managing AI risk, and 98% of surveyed organizations had made plans to increase their governance budgets.

Because MCP facilitates a one-way flow of trusted data into your existing tools, it’s governance-friendly. External products provide the data, but what is done with that data is entirely private to the user. Only users can make calls to the data, meaning that applications cannot see researchers’ prompts, enterprise’s internal data, or how AI is processing the information. This includes no “AI-to-AI” linkage, meaning that there is no need to be concerned that an enterprise’s internal data is leaked to MCP providers’ models. This is another reason why MCP integrations typically bypass the lengthy governance reviews required for new SaaS tools.

Discover Digital Science’s Dimensions & Altmetric MCPs

To maximize the potential of this innovation, Digital Science has designed Dimensions MCPs (Semantic Search and Analytics MCP) and Altmetric MCP to cater to the specific needs of research and data teams and provide verified, citable results so that users can discover more. 

Dimensions Semantic Search MCP translates plain-language questions into precise queries to run across one of the world’s most comprehensive databases, returning answers with full provenance. Through ontological concept resolution, the ontologically aware query, which drives the MCP, searches for concepts and corresponding synonyms. 

The Dimensions Analytics MCP provides a connected AI with a map of the research landscape, providing the links between relevant people, funding and organizations.

Altmetric shows where research is actually being read, shared, and acted on – in news media, policy documents, patents, and online conversations. Its MCP brings that attention data into your AI workflows so you can surface and report on societal impact at scale.

Key Benefits of Dimensions & Altmetric MCPs

Access to data from more than 430 million interconnected records, spanning publications, grants, patents, clinical trials, datasets, and policy documents, is just the beginning of what Dimensions and Altmetric offers research teams.

Despite the volume and complexity of the linked research data indexed, integrating Dimensions and Altmetric MCPs into a SaaS infrastructure requires neither complex code nor complex authentication. Because MCPs provide a standardization layer, users can move from setup to insights in minutes, not days, eliminating the need for heavy engineering resources.

The Altmetric MCP and two Dimensions MCPs are each designed to meet a specific need, creating a comprehensive research intelligence stack.  

Advanced content search with Dimensions Semantic Search MCP

  • Ontological concept resolution: Drug names, diseases, compounds are mapped to structured IDs across 40+ domains.
  • Co-occurrence discovery: Surface which drugs or compounds appear most with a given disease.
  • Multi-source search: Data from more than 430 million combined publications, patents, grants and clinical trials is considered in every query. 
  • Combined filters: Blend concept search with proximity, date, and author constraints.

Mapping the research ecosystem with Dimensions Analytics MCP

  • 430M+ linked records: Connect the dots between publications, grants, patents, clinical trials, datasets, policy documents, and even researchers and organizations.
  • Rich metadata: Find the most relevant data with robust metadata on research content records.
  • Natural Language & DSL: Query in plain English or use the full Dimensions Search Language.

Measuring impact with Altmetric MCP 

  • Attention Tracking: Monitor news, policy, social media, and patent references to research outputs in real time.
  • Filtering: Sort institutional research outputs by author, journal, and publication date.
  • Custom Analysis: Organize data by impact to your work – from publication-level details to aggregated metrics by therapeutic area, asset, key opinion leaders (KOL) impact, or company-level performance.

Dimensions Semantic Search MCP helps AI surface the content and evidence researchers need. It uses semantic technology to help teams query relevant material in publications, grants, patents, and clinical trials with just a key phrase or concept.

Dimensions Analytics MCP provides a linked view of the research ecosystem. It informs your AI workflow with research context from 430+ million linked publications, funding, researcher and organization profiles.

Altmetric MCP identifies domain experts and reveals the real-world impact of research and products via news, policy, social media, clinical guidelines, and patent monitoring.

Use Cases: From Discovery to Strategy

Here are some use case examples that demonstrate existing pain points teams face today, and how MCP, and particularly Dimensions and Altmetric MCPs, can help to solve their challenges.

Target Identification with Dimensions Semantic Search MCP

The problem: Researchers want to screen gene targets for a rare hereditary disease causing facial dysmorphism.

The strategy: Via the Dimensions Semantic Search MCP, the researchers are able to search in a harmonized way, using resolved concepts, across information from clinical trials, research publications, and patent data around the world and receive verified, citable results. The team is then able to visualize and generate reports with AI tools like GitHub and Microsoft Copilot, which can be integrated with Dimensions Semantic Search MCP. 

Cross-entity intelligence with Dimensions Analytics MCP

The problem: A biopharmaceutical company needs to find key opinion leaders (KOLs) and top researchers on a rare autoimmune condition for a steering committee.   

The strategy: Dimensions Analytics MCP can help the team identify key opinion leaders (KOLs) and accelerate strategic collaborations by mapping top researchers directly to their full scientific footprint. The Dimensions Analytics MCP connects you with deep, linked data on global publications, clinical trials, grant funding, intellectual property (IP), real-world impact metrics, and collaborative research networks. From broad therapeutic areas down to topics as granular as CRISPR-Cas12a off-target cleavage mechanics or AAV9 capsid engineering for crossing the blood-brain barrier in ALS, Dimensions Analytics MCP connects you to top researchers in any niche.

Research impact with Altmetric MCP 

The problem: A green chemistry firm publishes a landmark paper on a discovery they’ve made in catalyst technology, and wants to monitor the attention their breakthrough receives. 

The strategy: By linking an internal agent to the Altmetric MCP, the company can run automated reports of references to their research in real time.

Conclusion: Future-Proof Your AI Roadmap

At Digital Science, we’ve curated a range of products which help researchers push the boundary of discovery. And thanks to natural language discovery, the next big breakthrough in a research project could be unlocked via a simple prompt. In keeping with the Digital Science mission to democratize knowledge, the Dimensions Semantic Search MCP, Analytics MCP and Altmetric MCP are designed to make the research process as accessible and intuitive as possible, empowering non-technical team members to explore data from millions of publications and other research records via the built-in conversational interface or via the user’s agentic AI, linked via MCP. 

If you are interested in learning more about what AI research integration can achieve, contact us to see firsthand how the Dimensions and Altmetric MCPs can help.

The post Show your sources: building verifiable, citable AI agents with MCP appeared first on Digital Science.



from Digital Science https://ift.tt/TXdoacg

Dr. Zhuoyan Li | Medicine and Dentistry | Innovative Research Award


 Congratulations to Dr. Zhuoyan Li on receiving the Innovative Research Award at the Global Scholar Awards in Medicine and Dentistry. This distinguished honor recognizes your exceptional research achievements, pioneering innovations in medicine and dentistry, and unwavering dedication to advancing healthcare through scientific excellence, transformative discoveries, and meaningful contributions that create a lasting global impact.

Global Scholar Awards 🌟 Visit Our Website 🌐: globalscholarawards.com Nominate Now👍: https://globalscholarawards.com/doctor-awards-nobel-prize-scientists-award-nomination/?ecategory=Awards&rcategory=Awardee
Contact us ✉️: info@globalscholarawards.com

Get Connected Here: ================= Twitter : x.com/ScienceInventi1 Youtube : youtube.com/@nesinconferenceandawards4869 Pinterest : in.pinterest.com/scienceinventions/ Instagram : instagram.com/global_scholar_123 Linkedin : linkedin.com/in/global-scholar-awards-09664427b Blog : newscienceinventions2020.blogspot.com @WorldResearchAwards @GlobalScholarAwards #worldresearchawards #researchawards #researchexcellence #globalrecognition #academicawards #globalresearchawards #shorts #researchers #labtechnicians #awards #professors #teachers #lecturers

Featured Post

The next era of AI in drug discovery

As AI moves from summarizing papers to generating scientific claims, biopharma faces a hard question: how do you trust an answer with no wa...

Popular