Citation counts are attractive because they convert a difficult evaluative problem into a number. A publication accumulates citations; a researcher accumulates publications; an index combines the two; and committees gain an apparently objective basis for comparing people, projects, departments, and institutions. In research systems handling thousands of applications and promotion cases, that compression is operationally convenient.

The difficulty is that scientific impact is not a one-dimensional quantity. A highly cited paper may have changed a field, supplied a widely used method, become a standard reference, attracted criticism, or simply appeared in a rapidly publishing community. Conversely, a low-citation contribution may have produced a critical dataset, maintained software used in hundreds of analyses, prevented an unsafe engineering practice, improved a clinical protocol, trained a generation of researchers, established a difficult measurement capability, or solved a problem whose beneficiaries rarely publish papers.
Need help ? connect
This mismatch between what research systems can count and what research communities actually need has become a central issue in research assessment. The San Francisco Declaration on Research Assessment (DORA) explicitly warns against using journal-based metrics as surrogates for the quality of individual research contributions. The Leiden Manifesto similarly argues that quantitative evaluation should support, rather than replace, expert judgment and should account for differences among fields and research missions. More recent initiatives, including the Coalition for Advancing Research Assessment (CoARA), have pushed the argument further: assessment should recognize a broader range of outputs, practices, roles, and impacts.
The important question, therefore, is not whether citation metrics should disappear. Citations contain useful information. The question is what scientists should actually be rewarded for if the objective is a research system that produces reliable knowledge, useful capabilities, capable people, and durable public value.
What Citation Counts Measure—and What They Do Not
A citation is an observable scholarly event: one document refers to another. Aggregated across many papers, citations can reveal patterns of intellectual influence, field structure, knowledge diffusion, and scholarly attention. At sufficiently large scales and with careful normalization, bibliometrics can therefore be informative.
The problem begins when "citation" is silently translated into "quality," "importance," "impact," or "scientific merit." Those concepts overlap, but they are not equivalent.
Citation practices differ substantially between disciplines. Biomedical sciences, materials science, computer science, mathematics, engineering, and the humanities do not publish at the same rates, use the same document types, or accumulate citations on the same timescales. Even within one discipline, experimental papers, methods papers, datasets, reviews, theory papers, and software papers can have very different citation trajectories. A raw citation count therefore mixes the properties of the work with the properties of the communication system surrounding the work.
Time creates another distortion. Citation-based assessment systematically favors contributions old enough to accumulate a citation history. This is especially problematic when evaluating early-career researchers, emerging fields, or recent high-risk work. Short evaluation windows can reward topics with fast publication cycles while discounting research whose importance becomes visible only after years of validation, adoption, standardization, or technological maturation.
Citations also capture scholarly uptake more readily than many forms of external impact. A new finite-element formulation may be incorporated into commercial software without generating a proportional number of academic citations. An open-source package may become routine infrastructure. A geoscientific dataset may guide hazard planning. A validated materials-processing protocol may reduce industrial failure rates. A public-health study may affect a guideline. These are consequential outcomes, but their pathways do not necessarily return as references in indexed journal articles.
This is why responsible research assessment should treat citation data as evidence about one dimension of scholarly influence, not as a universal exchange rate for scientific value.
Why Citation-Centric Reward Systems Distort Research Behavior
Metrics do not merely observe research systems; once attached to careers and resources, they alter behavior within those systems. When a measure becomes a target, researchers rationally adapt to it. The issue is not that scientists become unethical by default. It is that incentives change the relative payoff of different activities.

A citation-centric environment makes some valuable work comparatively expensive. Maintaining community software, documenting a dataset, reproducing a disputed result, conducting a careful null study, building laboratory infrastructure, mentoring junior researchers, preparing standards, or collaborating deeply with industry can require substantial effort while producing fewer conventional publications. If hiring and promotion primarily reward papers and citations, researchers receive a clear signal that these activities should be secondary.
The same system can encourage fragmentation. A coherent research program may be divided into several publishable units because publication count is visible. Researchers may favor fashionable questions that offer larger citation markets. Review articles may become strategically attractive. Negative results and replication studies may be deprioritized because novelty is easier to publish and easier to signal.
The resulting problem is structural: individually rational behavior can produce a collectively weaker research ecosystem.
The Hong Kong Principles for assessing researchers were developed specifically to connect research assessment with research integrity. They emphasize responsible research practices, transparent reporting, open science, diversity of research contributions, and recognition of the full range of contributors. This is an important shift in perspective. A reward system should not merely identify visible outputs; it should reinforce the behaviors that make scientific outputs trustworthy.
Research Impact Is Better Represented as a Profile Than a Score
A more realistic model treats research impact as multidimensional. Conceptually, the impact of a researcher, team, or project can be represented as a vector
$
\mathbf{I} = (Q, R, O, C, K, T, S),
$
where $Q$ represents epistemic quality, $R$ research rigor and reproducibility, $O$ openness and reuse, $C$ contribution to collective research capability, $K$ knowledge and methodological contribution, $T$ translation into practice, and $S$ broader societal value.
The point of this representation is not to create seven new scores and combine them into another league table. In fact, blindly constructing
$
I_{\mathrm{total}} = \sum_i w_i I_i
$
with fixed weights $w_i$ risks reproducing the original problem in a more complicated form. Different research roles, career stages, disciplines, and institutional missions require different profiles. A theoretical physicist, research software engineer, clinical investigator, instrument scientist, and industry-facing mechanical engineer should not be expected to demonstrate impact through identical channels.
The useful principle is therefore structured multidimensionality: define the dimensions that matter, require evidence appropriate to each role, and use expert judgment to interpret the resulting profile in context.
| Dimension | What should be rewarded | Possible evidence |
|---|---|---|
| Epistemic quality | Important questions, sound reasoning, methodological appropriateness | Expert review, validation studies, robustness checks |
| Reproducibility and integrity | Transparent methods, traceable workflows, correction of errors | Reproducible packages, preregistration where appropriate, audit trails |
| Reusable outputs | Data, software, models, protocols, instruments | Repository records, persistent identifiers, documented reuse |
| Scholarly contribution | New concepts, methods, evidence, synthesis | Publications, citations used contextually, invited technical uptake |
| Team contribution | Collaboration, technical leadership, enabling work | Contribution statements, project evidence, infrastructure stewardship |
| Capacity building | Mentorship, supervision, training, community development | Trainee outcomes, curricula, workshops, technical documentation |
| Translation and societal impact | Influence on policy, standards, products, services, practice | Adoption evidence, standards contributions, case studies, implementation records |
Such a framework does not make evaluation easier in the sense of reducing it to a single number. It makes evaluation more valid.
Reward Epistemic Quality Before Visibility
The most important contribution of research is still the production of dependable knowledge. Any reformed assessment system should therefore begin with epistemic quality rather than downstream popularity.
Quality includes asking meaningful questions, choosing methods appropriate to those questions, reporting uncertainty, distinguishing exploratory from confirmatory analysis, testing alternative explanations, and avoiding stronger conclusions than the evidence supports. In computational work, it may include verification, validation, convergence studies, sensitivity analysis, uncertainty quantification, and clear specification of numerical assumptions. In experimental work, it can include calibration, controls, traceability, power considerations, and appropriate treatment of measurement error.
These activities are often poorly represented by headline metrics. A technically careful paper can be less citable than a broad claim. Yet the careful paper may be far more useful to experts deciding whether a result can be trusted.
Assessment therefore needs qualitative expert review, but "qualitative" should not mean impressionistic. Reviewers can be asked to identify the candidate's strongest contributions and examine the evidence supporting reliability, originality, and importance. The purpose is to move expert judgment from journal prestige and publication volume toward the substance of the work.
DORA's recommendations are especially relevant here: the scientific content of a paper should matter more than publication metrics or the identity of the journal in which it appeared.
Reward Reproducibility, Transparency, and Corrective Work
Science advances not only through new claims but through mechanisms that determine whether those claims survive scrutiny. Reproducibility work therefore creates scientific value even when it does not generate a novel headline result.
Researchers should receive credit for making analyses reproducible, publishing sufficient methodological detail, exposing assumptions, documenting computational environments, releasing validation cases, and correcting the scientific record when errors are found. Retractions and corrections should not automatically be treated as equivalent signals of failure. A transparent correction can demonstrate stronger research integrity than silent avoidance.
The relevant standard depends on the field. Reproducibility in a computational fluid dynamics study may require source code, mesh specifications, solver settings, boundary conditions, convergence criteria, and exact post-processing procedures. In clinical research, unrestricted data release may be impossible because of privacy and consent constraints, but reproducibility may still be strengthened through prespecified analysis plans, controlled-access data, clear protocol reporting, and executable analysis code.
The objective is not "open everything." It is to reward researchers for making their work as inspectable, testable, and reusable as legal, ethical, technical, and security constraints permit.
The UNESCO Recommendation on Open Science connects open scientific knowledge, infrastructure, engagement, and transparent evaluation with a broader vision of high-quality research. Similarly, the NIH Data Management and Sharing Policy institutionalizes the expectation that scientific data should be managed and shared appropriately, recognizing that access and reuse can accelerate discovery and validation.
If institutions claim to value reproducibility while promotions ignore the labor required to achieve it, policy and incentives remain misaligned.
Reward Research Software, Data, Methods, and Infrastructure as First-Class Outputs
Modern science is increasingly built on research objects that are not conventional papers. Simulation codes, machine-learning pipelines, reference datasets, instrument designs, ontologies, laboratory protocols, benchmark suites, computational workflows, and curated databases can be foundational to entire communities.
Yet authorship and citation systems historically concentrated credit around publications. The consequence is familiar in computational science: a researcher may spend years building robust software that enables dozens of projects, while career evaluation treats the software as supporting work and rewards the papers produced by its users.
This is not simply unfair to developers. It creates a maintenance problem. If software quality, testing, documentation, versioning, and long-term stewardship are not rewarded, research institutions encourage the creation of fragile infrastructure and then become dependent on it.
The FORCE11 Software Citation Principles argue that software should be treated as a legitimate research product and cited on a comparable basis to other scholarly outputs. The same logic applies to carefully curated datasets. Persistent identifiers, versioned releases, structured metadata, documentation, validation tests, and explicit reuse records make these outputs assessable.
A strong promotion or funding dossier should therefore be able to say not only "this researcher published 30 papers," but also "this researcher created the solver used by five collaborating laboratories," "this dataset became the reference benchmark for a measurement problem," or "this protocol eliminated a recurrent source of experimental variability."
Those are research contributions in the direct sense of the term.
Reward What People Actually Contributed
Traditional authorship compresses complex collaboration into an ordered list of names. That representation is often too coarse for modern team science.
A paper may depend on conceptual design, data curation, software, formal analysis, laboratory investigation, visualization, project administration, supervision, validation, and acquisition of resources. The CRediT Contributor Role Taxonomy provides 14 standardized contributor roles precisely because authorship position alone cannot reliably describe this distribution of labor.
Assessment systems should use contribution statements as evidence, particularly for large collaborations. This is important for researchers whose work is technically central but less visible in conventional authorship hierarchies: data scientists, research software engineers, facility scientists, statisticians, instrumentation specialists, project managers, and technical staff.
The aim is not to create a bureaucratic ledger in which every hour is attributed. Rather, evaluators need enough information to answer a basic question: what did this person make possible?
That question is especially important when assessing leadership. Scientific leadership is not synonymous with being the corresponding author. A researcher may lead by designing an architecture, establishing a measurement capability, coordinating a consortium, resolving a methodological bottleneck, creating quality-control processes, or mentoring others to independence.
Reward Replication, Negative Results, and Methodological Maintenance
Research systems frequently say that replication is essential while rewarding novelty more strongly. This creates a predictable underproduction of confirmatory work.
A well-designed replication can determine whether a result generalizes across laboratories, populations, instruments, computational settings, or environmental conditions. A rigorous negative result can close an unproductive line of inquiry or reveal that an apparent effect depends on a narrow parameter range. A benchmark study can expose numerical instability in a popular method. A maintenance release can prevent errors across an entire software user base.
These outputs often have high system-level value because they reduce uncertainty and prevent waste. Their importance should therefore be assessed in relation to the decisions they enable, not only the novelty of their conclusions.
A mature reward framework should make room for researchers who improve the reliability of a field, not only those who expand its catalogue of claims.
Reward Mentorship, Training, and the Creation of Research Capability
Scientific output is partly the product of accumulated human capability. Researchers build that capability when they train students, mentor colleagues, establish good laboratory practices, create technical courses, document specialist methods, and help teams develop new competencies.
These activities can have long time horizons. A principal investigator who creates an environment in which doctoral researchers learn rigorous experimental design may influence dozens of future projects. A senior engineer who develops reusable training around finite-element verification can improve the quality of simulation work across an organization. A scientist who mentors early-career researchers into independence contributes to research capacity even when the resulting papers are published after the formal evaluation window.
Narrative CV approaches attempt to make such contributions visible. The Royal Society's Résumé for Researchers was designed to capture broader contributions including development of others and contributions to the research community. The UK Research and Innovation Résumé for Research and Innovation (R4RI) similarly provides a framework for evidencing a wider range of skills and experiences than a conventional publication-centered CV.
Mentorship should not be rewarded merely because it occurred. Evidence still matters. Useful indicators may include trainee progression, documented supervision practices, development of shared technical resources, contributions to training programs, and examples showing that mentees gained autonomy or new capability.
Reward Translation Without Requiring Every Scientist to Become an Entrepreneur
"Impact" is sometimes narrowed too aggressively toward commercialization. Patents, licenses, spinouts, industrial contracts, and revenue can be important, particularly in engineering and applied science, but they represent only some pathways from knowledge to use.
Translation can include incorporation into standards, regulatory guidance, clinical pathways, public policy, open-source engineering tools, industrial design rules, environmental monitoring, educational practice, or professional workflows. It can also involve reducing uncertainty for decision-makers, even when there is no product.
The critical requirement is evidence of a plausible pathway from research to changed practice or capability. A policy citation alone does not prove policy impact. A patent alone does not prove technological impact. A software download does not prove meaningful use. Strong assessment examines the chain connecting contribution to adoption and outcome.
At the same time, basic researchers should not be penalized for lacking immediate societal or commercial outcomes. Some scientific contributions have long and uncertain translation pathways. The correct question is whether the claimed type of impact is appropriate to the mission and maturity of the work.
This contextual principle is central to responsible assessment. It prevents a theoretical mathematician from being judged by industrial revenue and prevents an applied engineering program from being judged solely by journal citations.
Reward Open Science When Openness Produces Scientific Utility
Open science is valuable because it can increase inspectability, accessibility, reuse, and participation. But assessment should distinguish substantive openness from procedural box-ticking.
Uploading undocumented files to a repository shortly before evaluation has limited value. A genuinely reusable dataset needs appropriate metadata, provenance, documentation, licensing, and where necessary access controls. Reusable software requires installation instructions, dependency management, tests, versioning, examples, and maintenance practices. An open protocol should be sufficiently precise for competent practitioners to implement it.
The appropriate reward is therefore not "number of open outputs." It is the quality and utility of open research practices.
Evidence of reuse can be especially informative: independent software adoption, dataset reuse, forks and contributions, integration into workflows, replication enabled by shared materials, or community extensions. These signals should still be interpreted carefully, because highly specialized resources may be extremely important to small communities.
Reward Contributions to Research Culture and Collective Reliability
Research quality is influenced by local culture: how teams discuss uncertainty, handle errors, share credit, respond to failed experiments, supervise trainees, document workflows, and manage pressure. These practices rarely appear in bibliometric databases, but they affect the reliability and sustainability of scientific work.
Institutions should therefore recognize contributions such as establishing reproducible workflow standards, improving laboratory safety, developing fair authorship practices, supporting research integrity, building shared infrastructure, and creating inclusive mechanisms for technical decision-making.
This direction is increasingly visible in institutional assessment. For example, the UK's Research Excellence Framework 2029 Strategy, People and Research Environment guidance explicitly treats research strategy, people, and environment as part of the assessment architecture. The exact design of national systems will differ, but the underlying point is general: excellent outputs are partly produced by excellent research environments.
The risk, however, is replacing one vague proxy with another. "Good citizenship" should not become a subjective popularity criterion. Claims about culture and leadership should be supported by concrete activities, changes, and outcomes.
A Better Assessment Architecture
Replacing citation counts requires more than adding new metrics to an old scorecard. The assessment process itself needs redesign.
Start With the Purpose of the Decision
Hiring, promotion, fellowship selection, grant review, institutional evaluation, and research-prize decisions have different objectives. The evidence requested should reflect the decision.
A grant panel may need to evaluate whether a team has the capability to execute a proposed program. A promotion committee may need to evaluate sustained contribution and trajectory. An institution-level review may need to assess collective research quality, culture, infrastructure, and external impact.
The first design principle is therefore to define what success means before choosing indicators.
Ask for a Small Number of Significant Contributions
Long publication lists encourage counting. A stronger approach asks researchers to identify a limited number of important contributions and explain their significance, their role, the evidence supporting impact, and the context in which the work should be interpreted.
This changes the review task from "scan for prestige signals" to "evaluate claims and evidence."
A contribution could be a paper, but it could also be software, a dataset, an instrument, a methodological framework, a standard, a translational program, or a body of mentorship and infrastructure work. The evidentiary burden should be comparable: the candidate must show why the contribution matters.
Separate Evidence From Interpretation
Assessment dossiers should distinguish between observations and conclusions.
For example:
- Observation: a software package has 2,000 documented dependent projects.
- Interpretation: the software has become important research infrastructure in its field.
- Observation: a method is cited in a regulatory technical document.
- Interpretation: the research influenced regulatory practice.
- Observation: three independent groups reproduced a result using the released workflow.
- Interpretation: the work demonstrated unusually strong reproducibility.
The observation can often be verified. The interpretation requires expert judgment. Keeping these layers distinct reduces the temptation to treat every indicator as self-explanatory.
Use Metrics Diagnostically, Not Mechanically
Citation counts, field-normalized citation indicators, software downloads, repository reuse, patent citations, policy mentions, clinical guideline citations, standards references, invited talks, and collaboration networks can all be informative. None should become a universal threshold.
The Leiden Manifesto remains useful precisely because it frames metrics as contextual evidence and emphasizes field variation, transparency, and scrutiny of indicators.
Metrics are strongest when they answer a specific question. "Has this method diffused beyond the originating group?" is a meaningful question. "Is an h-index of 32 good enough for promotion?" is usually not, unless an institution is willing to defend why that threshold validly represents the qualities required for the role.
Calibrate Peer Review
Moving away from simple metrics increases dependence on judgment, and judgment has its own failure modes: prestige bias, institutional bias, familiarity effects, conservatism, disciplinary narrowness, and inconsistent standards.
Responsible qualitative assessment therefore needs structure. Reviewers should receive explicit criteria, distinguish contribution from reputation, disclose conflicts, justify conclusions with evidence, and calibrate interpretations across cases. Committees should be prepared to explain why a candidate was strong or weak without relying on journal brands as shorthand.
The goal is not metric-free assessment. It is auditable expert assessment.
What This Looks Like in Different Research Domains
The value of multidimensional assessment becomes clearer when applied to concrete cases.
Computational Science and Engineering
Consider a researcher who develops an open multiphysics solver. The conventional output may include several papers, but much of the scientific contribution lies elsewhere: numerical verification, benchmark problems, documentation, test coverage, solver stability, reproducible examples, user support, and integration by other groups.
A meaningful evaluation would examine software quality, independent use, documented validation, contributions to standards or benchmark practices, and the research enabled by the tool. Citation counts can supplement that picture but should not define it.
Experimental Science
An experimental researcher may create a measurement platform with exceptional calibration stability and release a reference dataset used to compare techniques across laboratories. The instrument paper itself may receive moderate citations, while the platform changes what the field can measure.
Assessment should recognize creation of measurement capability, uncertainty characterization, data quality, reproducibility across operators, and the downstream work enabled by the facility or method.
Biomedical and Clinical Research
A clinical investigator may contribute to a negative multicenter trial that prevents adoption of an ineffective intervention. Scientifically and socially, that result can be highly valuable even if it is less celebrated than a successful therapeutic result.
Relevant evidence includes study design quality, protocol adherence, transparent reporting, data stewardship, contribution to guidelines, and implications for clinical decisions. Rewarding only positive novelty would create precisely the wrong incentive.
Applied Engineering and Industry Collaboration
An engineering team may develop a structural assessment method that is incorporated into a company's design workflow and reduces repeated prototype testing. Some details may remain confidential, limiting publication.
Evaluation can still examine technical validation, documented adoption, changes to engineering decisions, transfer of capability, standards participation, and independently verifiable outcomes. A publication-only framework would systematically undervalue this work.
What Institutions Should Stop Rewarding by Default
Reform also requires subtraction. Institutions should stop treating several common signals as automatic evidence of excellence.
Journal prestige should not substitute for reading the work. Raw publication counts should not substitute for understanding contribution. Total citations should not be compared across unrelated fields without context. Authorship position should not be treated as a universal measure of contribution. Grant income should not be confused with scientific impact; it is primarily an input that enables research. Patent counts should not be equated with innovation without evidence of use. Social-media attention should not be treated as equivalent to scientific or societal influence.
Most importantly, institutions should avoid building a giant dashboard and assuming that more indicators automatically produce fairer assessment. A sufficiently complicated metric system can be harder to understand, easier to game, and less accountable than the system it replaces.
CoARA's Agreement on Reforming Research Assessment is useful here because it places qualitative assessment and responsible use of quantitative indicators within a broader institutional reform agenda rather than presenting a single replacement metric.
A Practical Reward Framework for Scientists
A workable evaluation system can be built around five questions.
First, what important knowledge or capability did the researcher help create? This establishes the substantive contribution.
Second, how trustworthy is the work? This brings rigor, integrity, validation, uncertainty, and reproducibility into the center of evaluation.
Third, what did the researcher personally contribute? This corrects for the inadequacy of publication counts and author order in team science.
Fourth, who or what became more capable because of the work? This captures software, datasets, infrastructure, mentorship, methods, community resources, and institutional capability.
Fifth, what evidence exists that the contribution influenced scholarship, practice, technology, policy, or society? Citations can answer part of this question, but only one part.
This framework also creates space for legitimate specialization. Not every scientist needs to excel in every category. A research software engineer may have extraordinary infrastructure and reuse impact. A theoretician may transform conceptual understanding. A laboratory leader may create unique measurement capability. A clinician-scientist may change practice. A senior researcher may have exceptional mentorship and field-building impact.
Evaluation should identify excellence in the context of role and mission rather than forcing all careers toward the same publication-maximizing template.
If you're working on related challenges in this area and would find guidance helpful, feel free to reach out: CONTACT US.
Conclusion
Citation counts are not meaningless. They are simply too narrow to carry the weight placed on them. They measure a form of scholarly attention, imperfectly and with strong dependence on field, time, document type, and community practice. Used carefully, they can contribute to research assessment. Used as a proxy for scientific worth, they distort it.
Scientists should be rewarded for producing reliable knowledge; making methods, data, and software reusable; strengthening reproducibility and research integrity; contributing meaningfully to teams; mentoring people; building research capability; correcting the scientific record; translating knowledge into useful practice; and creating societal, technological, clinical, environmental, or policy value where such outcomes are appropriate to the work.
The central design principle is not to replace one number with seven new numbers. Research impact should be evaluated as an evidence-backed profile interpreted in context. Quantitative indicators can inform that profile, but they should not determine it mechanically.
A research system ultimately gets more of what it rewards. If institutions reward citation accumulation, researchers will optimize for visible scholarly output. If they reward rigor, reuse, contribution, capability, and demonstrated influence, they create stronger incentives for the forms of work on which durable scientific progress actually depends.
Interested in collaborating on academic research ? feel free to get in touch 🙂.
Check out YouTube channel, published research
you can contact us (bkacademy.in@gmail.com)
Interested to Learn Engineering modelling Check our Courses 🙂