Keywords are a small component of a research manuscript, but they perform several important functions. They help journal submission systems classify a paper, support editor and reviewer assignment, contribute to bibliographic indexing, and improve the probability that a relevant reader will retrieve the article through a database query. In Q1 journals, keywords do not compensate for weak science, poor journal fit, or an unclear abstract. However, poorly chosen keywords can reduce discoverability, create ambiguity about the paper’s scope, and occasionally route the manuscript toward an unsuitable editor or reviewer pool.
Need help ? connect
Choosing keywords should be treated as an information-retrieval problem rather than a last-minute administrative task. The goal is not to list every term appearing in the manuscript. It is to construct a concise representation of the article that aligns the language of the study with the language used by the target research community, indexing systems, and prospective readers.

A strong keyword set combines the central research object, the principal phenomenon or problem, the methodological approach, the relevant application or context, and one or more established disciplinary terms. The balance depends on the journal, field, article type, and permitted number of keywords. This article develops a rigorous procedure for selecting, testing, and refining author keywords for research articles intended for Q1 journals.
Why keywords matter in Q1 journal publishing
Q1 is a journal-ranking designation rather than a universal editorial standard, and a journal can occupy different quartiles across subject categories. Nevertheless, upper-quartile journals usually operate within dense, competitive literatures, making accurate classification and discoverability important. Keywords influence both submission and post-publication workflows.
First, many editorial management systems use manuscript keywords, subject classifications, or both to assist with assigning handling editors and identifying reviewer candidates. The precise algorithms are usually proprietary, and keywords are only one input among many. Yet a vague or misleading keyword set can signal the wrong disciplinary emphasis. For example, an article on physics-informed neural networks for turbulent heat transfer should not be represented only by broad terms such as “artificial intelligence,” “simulation,” and “engineering.” Those terms fail to identify the governing-equation framework, transport phenomenon, and specific modelling task.
Second, bibliographic platforms expose author keywords as searchable metadata. Scopus distinguishes author keywords from indexed keywords, while also searching titles and abstracts. Keywords should therefore be designed as part of a coherent metadata package rather than as an isolated list.
Third, readers use terminology unevenly. A single concept may be searched through a formal name, an abbreviation, an older term, a neighbouring disciplinary term, or a specific application phrase. Keywords provide an opportunity to cover important terminology that cannot be included naturally in the title. Springer Nature’s guidance on titles, abstracts, and keywords emphasizes their joint role in helping readers find research and judge its relevance. This relationship is central: keyword selection should complement the title and abstract, not merely copy them.

Finally, keywords shape the paper’s long-term semantic position. Idiosyncratic terminology reduces discoverability, while fashionable but peripheral terms attract irrelevant searches and weaken precision.
Author keywords are not the same as search-engine keywords
Researchers sometimes approach manuscript keywords using commercial search-engine optimization practices. The analogy is only partially useful. Scholarly retrieval differs from general web search in important ways.
Academic databases search structured fields, index controlled vocabularies, normalize bibliographic records, link citations, and apply field-specific ranking procedures. Journal websites may also be indexed by general search engines, but the author’s primary responsibility is accurate scientific description. The objective is not to maximize traffic from any possible query. It is to improve retrieval by the correct audience without distorting the content.
A manuscript keyword should therefore satisfy three requirements:
- Topical validity: the term must describe a substantial component of the study.
- Community validity: the term should be recognizable to researchers in the relevant field.
- Retrieval value: the term should add useful search coverage beyond what is already obvious from the title.
A keyword that is popular but peripheral fails the first requirement. A private laboratory abbreviation fails the second. A term that repeats a distinctive title phrase without adding any new retrieval route may contribute little to the third.
Begin with the target journal’s instructions
The first step is not brainstorming. It is reading the target journal’s current guide for authors and submission form requirements. Journals differ in the number, format, capitalization, and permitted structure of keywords. Some request three to six terms; others allow a larger range. Some strongly prefer single words or short phrases. Some require authors to choose from a predefined taxonomy during submission. Others ask for both free-text keywords and subject classifications.
These requirements override general advice. The National Library of Medicine’s guidance for authors selecting MeSH descriptors explicitly advises authors to consult the instructions of the specific journal because requirements vary. Similarly, IEEE Access submission guidance specifies a minimum and maximum number of manuscript keywords and notes that accurate keywords support the matching of a relevant associate editor.
Before selecting terms, record the following constraints:
| Constraint | Question to verify |
|---|---|
| Number | How many keywords are required or permitted? |
| Format | Are phrases allowed, or are single terms preferred? |
| Source | Must terms come from a journal taxonomy or controlled vocabulary? |
| Language | Must keywords be supplied only in English or in multiple languages? |
| Duplication | Does the journal discourage repeating words from the title? |
| Submission fields | Are author keywords separate from classifications, topics, or index terms? |
Do not infer these rules from another journal published by the same company. Requirements may differ across titles, article types, and editorial systems.
Define the manuscript’s semantic core
A reliable keyword set begins with a structured representation of what the paper actually contributes. Before generating terms, reduce the manuscript to five semantic dimensions:
- Research object: What system, material, population, dataset, device, process, or theory is being studied?
- Problem or phenomenon: What behaviour, limitation, mechanism, response, or scientific question is addressed?
- Method: What experimental, computational, analytical, statistical, or theoretical approach is central?
- Context or application: In what environment, scale, industry, clinical setting, or operating regime does the study matter?
- Distinctive contribution: What differentiates the paper from nearby literature?
This decomposition prevents an overly broad list. Consider a hypothetical manuscript titled “Physics-Informed Neural Network Reconstruction of Temperature Fields in Sparse-Sensor Convection Experiments.” Its semantic core might be represented as follows:
| Dimension | Candidate concept |
|---|---|
| Research object | Temperature field |
| Problem | Sparse-sensor reconstruction |
| Method | Physics-informed neural network |
| Governing context | Convective heat transfer |
| Distinctive feature | Experimental data assimilation |
The final keywords should not necessarily reproduce all five entries verbatim. Instead, these concepts form the candidate pool from which terms are selected after considering journal scope, disciplinary vocabulary, and title overlap.
Build a candidate keyword inventory
The most effective workflow separates candidate generation from final selection. Generate a relatively broad inventory first, then evaluate each term using explicit criteria.
Extract terms from the manuscript
Review the title, abstract, objectives, methods, principal results, figure captions, and conclusion. Record recurring technical nouns and established noun phrases. Give particular attention to terms that appear in the research question, define the method, or distinguish the contribution.
Term-frequency counts can assist, but frequency is not importance. Generic words such as “model,” “analysis,” and “performance” may dominate the manuscript while adding little retrieval specificity; a distinctive material or algorithm may appear less often but remain essential.
Examine recent papers in the target journal
Search the target journal for recent articles that are close to the manuscript in topic and article type. Inspect their author keywords, titles, and abstracts. The purpose is not to copy keywords mechanically. It is to identify the journal’s active vocabulary and the level of specificity commonly used.
For example, a broad engineering journal may use “finite element method,” while a specialized computational mechanics journal may favour more precise terms such as “phase-field fracture,” “isogeometric analysis,” or “nonlinear model updating.” The appropriate granularity depends on the readership.
Review roughly 10–20 closely related papers from the target journal and adjacent high-quality journals. Recurring terms indicate vocabulary stability, but frequency alone does not establish suitability; a common field term may still fail to identify the paper’s contribution.
Inspect controlled vocabularies and disciplinary taxonomies
Controlled vocabularies reduce terminological variation through standardized descriptors. In biomedicine, Medical Subject Headings, or MeSH, is the U.S. National Library of Medicine vocabulary used for PubMed and MEDLINE indexing. Authors can inspect preferred descriptors, entry terms, and broader or narrower concepts.
Engineering, computing, chemistry, geology, psychology, agriculture, and other disciplines use their own taxonomies or thesauri. When the submission system supplies a controlled list, select from it directly. Because standardized descriptors may be too broad for an emerging method or specific application, combine them with accurate free-text phrases when the journal permits.
Identify synonyms, abbreviations, and spelling variants
For each core concept, list plausible variants:
- full term and accepted abbreviation;
- singular and plural forms where meaning changes;
- British and American spellings;
- older and newer terminology;
- method name and algorithm family;
- material class and specific material;
- clinical or common name and standardized descriptor.
Do not include every variant. Select the form with the strongest disciplinary and retrieval value. “Finite element analysis” and “finite element method,” for example, are related but not always interchangeable, while database normalization may already handle spelling variants such as “fibre” and “fiber.”
Evaluate candidates using a scoring model
Keyword selection can be made more consistent by scoring candidates against a small set of criteria. A practical model is:
$
S_i = w_r R_i + w_s P_i + w_c C_i + w_j J_i + w_d D_i - w_a A_i,
$
where $S_i$ is the score for candidate keyword $i$; $R_i$ is relevance to the central contribution; $P_i$ is specificity; $C_i$ is consistency with community terminology; $J_i$ is alignment with the target journal; $D_i$ is incremental discoverability beyond the title; and $A_i$ is ambiguity or risk of attracting irrelevant searches. The coefficients $w_r$, $w_s$, $w_c$, $w_j$, $w_d$, and $w_a$ are weights reflecting the importance assigned to each factor.
Authors do not need to calculate a formal numerical score. The equation is a disciplined way to express the trade-off. Relevance should normally receive the largest weight. A term with high search volume but weak relevance should not survive selection.

A simple qualitative matrix is often sufficient:
| Candidate | Central relevance | Specificity | Community use | Adds title coverage | Ambiguity | Decision |
|---|---|---|---|---|---|---|
| Artificial intelligence | Medium | Low | High | Medium | High | Reject or replace |
| Physics-informed neural networks | High | High | High | High | Low | Retain |
| Heat transfer | High | Medium | High | Medium | Medium | Retain if scope permits |
| Sparse sensor data | High | High | Medium | High | Low | Retain |
| Numerical simulation | Medium | Low | High | Low | High | Reject |
This process prevents the keyword set from becoming a collection of broad, prestigious-sounding terms.
Balance breadth and specificity
A keyword set should support retrieval at more than one conceptual level. If every term is extremely broad, the paper disappears into a large literature. If every term is extremely narrow, researchers using broader queries may never encounter it.
A useful architecture is a layered set:
- One field-level term that places the article in the principal discipline or subdiscipline.
- One or two object or phenomenon terms that describe what is being studied.
- One or two methodological terms that identify the central approach.
- One application or distinguishing term that captures the specific context or contribution.
Suppose a paper reports a graph neural network surrogate for transient fluid flow on unstructured meshes. A weak keyword set might be:
“Artificial intelligence; modelling; simulation; prediction; fluid mechanics.”
A stronger set might be:
“Graph neural networks; mesh-based learning; transient flow; surrogate modelling; computational fluid dynamics; unstructured meshes.”
The improved set contains a recognizable field term, a method family, a representation, a physical problem, and a computational purpose. It is also less ambiguous.
Coordinate keywords with the title and abstract
The title, abstract, and keyword list should behave as a coherent retrieval system. They should not be identical, but they should not contradict one another.
A useful principle is controlled redundancy. The manuscript’s most important concept should appear in the title or abstract and may also appear as a keyword when the journal permits or expects repetition. However, every keyword need not repeat title wording. Elsevier’s manuscript-preparation guidance recommends avoiding unnecessary duplication of title words and using related or method-specific terms to complement the main topic. This is sensible when the title already contains a distinctive phrase.
For example, if the title already includes “silicon photonic neural network accelerator,” the keyword list may add “wavelength-division multiplexing,” “optical matrix multiplication,” “inference hardware,” and “energy efficiency,” assuming these are central to the study. Repeating all title words individually would use limited keyword slots without expanding semantic coverage.
At the same time, avoid forcing synonyms that are not standard in the field simply to eliminate repetition. Exact duplication is less harmful than introducing an inaccurate substitute. Journal rules and disciplinary norms should determine the balance.
Choose noun phrases, not descriptive sentences
Keywords usually function best as concise nouns or noun phrases. They should identify concepts, not summarize findings. Avoid complete clauses such as “improves the accuracy of prediction” or “a new method for detecting defects.” Convert these into stable concepts such as “prediction accuracy,” “defect detection,” or the specific method name.
Good keyword phrases are generally:
- semantically complete;
- recognizable without the surrounding sentence;
- free of unnecessary articles and prepositions;
- short enough to match likely queries;
- precise enough to distinguish the study.
Overly long phrases often become fragile because readers may search the same concept using a shorter formulation. For example, “deep-learning-based automated crack-detection method for concrete bridges” is better decomposed into “deep learning,” “crack detection,” “concrete bridges,” and possibly the specific imaging method.
However, indiscriminate decomposition can also reduce meaning. “Physics-informed neural networks” should normally remain a phrase rather than being split into “physics,” “informed,” and “neural networks.” Multiword expressions should be preserved when they represent a recognized technical concept.
Avoid common keyword-selection errors
Selecting terms that are too broad
Terms such as “science,” “engineering,” “optimization,” “performance,” “analysis,” “model,” and “technology” are rarely useful by themselves. Their document frequency is so high that they provide little discriminatory value. Broad terms may be retained only when they define an established field and are combined with more specific terms.
Using fashionable but peripheral terminology
Including “artificial intelligence,” “digital twin,” “sustainability,” “Industry 4.0,” or another high-visibility term is inappropriate unless the concept is materially developed in the paper. Editorial reviewers often notice when keywords overstate the scope. This can create a credibility problem before the technical evaluation begins.
Repeating near-synonyms
A limited keyword allowance should not be consumed by terms such as “finite element method,” “finite element analysis,” and “FEM” unless the journal’s retrieval system or field conventions justify including multiple forms. In most cases, choose the preferred expression and use the abbreviation naturally in the abstract.
Using undefined abbreviations
An abbreviation may be suitable when it is universally recognized in the target field, such as CFD, DNA, or MRI, although journal conventions vary. A local abbreviation, newly coined acronym, or ambiguous initialism should not normally be used alone. “GNN,” for example, can be useful in a machine-learning context but is less informative than “graph neural networks” as an author keyword.
Including unsupported terms
A reader should reasonably expect substantial discussion, evidence, or results for every keyword. Remove any term that cannot be traced to a meaningful part of the manuscript.
Ignoring the article type
Keywords for a review article may need to represent a broader conceptual landscape than those for a narrowly focused experimental paper. A methods paper should emphasize the methodological innovation and validation domain. A dataset paper should identify the data type, domain, acquisition or generation method, and intended use. A clinical trial may require population, intervention, condition, and study design terms.
Copying keywords from a single paper
Even an influential article may use older terminology or follow another journal’s conventions. Base keyword design on a representative literature sample.
Use literature databases as testing environments
Candidate keywords should be tested in the databases that the target audience actually uses. The objective is to inspect the neighbourhood of literature retrieved by each term.
For every candidate, run a search and examine the first 20–50 results. Ask:
- Are most results from the intended field?
- Do the retrieved articles resemble the manuscript in topic or method?
- Is the term dominated by another discipline?
- Does the query retrieve an unmanageably broad literature?
- Does a more standardized or more specific variant perform better?
- Are recent papers using a different preferred term?
This is not a search-volume competition. A highly specific keyword may retrieve fewer papers but show much stronger topical precision. In information retrieval, this trade-off is commonly expressed through precision and recall:
$
\text{Precision} = \frac{\text{relevant documents retrieved}}{\text{all documents retrieved}},
$
$
\text{Recall} = \frac{\text{relevant documents retrieved}}{\text{all relevant documents in the collection}}.
$
A broad term tends to improve recall but reduce precision. A narrow term tends to improve precision but may reduce recall. A balanced keyword set uses multiple terms to cover both objectives without becoming redundant.
Search testing is especially valuable for interdisciplinary manuscripts. A phrase may have different meanings in materials science, computer science, economics, or medicine. For instance, “resilience,” “domain adaptation,” “fatigue,” and “translation” all have multiple disciplinary interpretations. Adding an appropriate qualifier can prevent the paper from being semantically misplaced.
Distinguish author keywords from controlled index terms
Author keywords are chosen by the authors and submitted with the manuscript. Indexed terms may be assigned later by database algorithms, professional indexers, or publisher systems. The two sets can overlap, but they are not interchangeable.
This distinction has practical consequences. Authors should not assume that selecting a broad author keyword guarantees the same controlled descriptor will be assigned after publication. Nor should they assume that database indexing makes author keywords irrelevant. Author keywords remain part of the published metadata in many systems and may be searchable as a distinct field. IEEE Xplore, for example, provides an “Author Keywords” field in its command-search documentation.
In biomedical publishing, an author may select MeSH-compatible terms, but final MEDLINE indexing is determined through NLM indexing processes. The most defensible approach is to use standardized descriptors where they fit while preserving necessary free-text specificity for emerging concepts.
A practical step-by-step workflow
The following workflow can be incorporated into manuscript preparation before submission.
Step 1: Freeze the scientific scope
Do not finalize keywords while the manuscript’s central claim is still changing. Complete the title, abstract, primary figures, and conclusion first. Keyword selection should reflect the final version of the contribution.
Step 2: Record journal constraints
Extract the keyword count, formatting rules, controlled lists, and duplication policy from the journal’s current instructions and submission interface.
Step 3: Write a one-sentence contribution statement
Use a technical sentence containing the object, problem, method, and principal contribution. For example:
“This study develops a multi-fidelity Gaussian-process surrogate for uncertainty quantification in nonlinear finite element models of composite pressure vessels.”
The major noun phrases become the initial candidate concepts.
Step 4: Generate 12–20 candidates
Collect candidates from the manuscript, target-journal papers, controlled vocabularies, and accepted synonyms. At this stage, retain both broad and narrow possibilities.
Step 5: Remove invalid candidates
Delete terms that are generic, peripheral, promotional, ambiguous, unsupported, or duplicative. Also remove terms that violate the journal’s formatting requirements.
Step 6: Test the remaining terms
Search the target journal, Scopus, Web of Science, PubMed, IEEE Xplore, Google Scholar, or the principal database used in the field. Compare retrieval precision and disciplinary fit.
Step 7: Construct a balanced set
Select terms that collectively cover the research object, problem, method, and context. Avoid allowing one dimension, usually the method, to occupy every slot.
Step 8: Check title–abstract–keyword consistency
Confirm that the selected terms are supported by the abstract and manuscript. Add an important term to the abstract when scientifically appropriate, but do not force unnatural repetition.
Step 9: Ask a domain expert to perform a retrieval test
Show a colleague only the title, abstract, and keywords. Ask what they think the paper is about and which literature they would expect it to cite. Misinterpretation at this stage signals a metadata problem.
Step 10: Revalidate during submission
Some systems present a controlled list after the manuscript files are uploaded. Reconcile the free-text keywords with the available classifications. Do not select the closest-looking category without checking its actual disciplinary meaning.
Worked examples across research domains
Example 1: Computational mechanics
Study: A phase-field finite element model for fatigue crack growth in additively manufactured titanium alloy components.
Weak keywords: simulation; fatigue; metal; engineering; modelling.
Improved keywords: phase-field fracture; fatigue crack growth; additive manufacturing; titanium alloys; finite element method; damage mechanics.
The improved set identifies the fracture formulation, failure process, manufacturing route, material class, numerical method, and theoretical context. “Simulation” and “engineering” add little because they are too broad.
Example 2: Photonics
Study: Inverse design of a silicon photonic wavelength demultiplexer using an adjoint optimization method.
Weak keywords: optics; design; silicon; optimization; device.
Improved keywords: silicon photonics; wavelength demultiplexing; inverse design; adjoint optimization; nanophotonic devices; integrated optics.
Here, “silicon” becomes the established field phrase “silicon photonics,” and “optimization” is refined to the actual method. The keywords also include both the device function and the broader integration context.
Example 3: Biomedical research
Study: A prospective study of MRI-derived radiomic markers for predicting treatment response in glioblastoma.
Weak keywords: cancer; imaging; prediction; brain; treatment.
Improved keywords: glioblastoma; magnetic resonance imaging; radiomics; treatment response; predictive biomarkers; prospective studies.
A biomedical author should additionally inspect MeSH to determine the preferred descriptors and confirm the journal’s requirements. The improved set avoids the very broad term “cancer” and identifies the tumour type, modality, analytical approach, outcome, biomarker role, and study design.
Example 4: Machine learning for engineering
Study: A graph neural network trained on finite element meshes to predict stress fields under variable boundary conditions.
Weak keywords: AI; neural network; stress; data; prediction.
Improved keywords: graph neural networks; mesh-based learning; stress-field prediction; finite element meshes; surrogate modelling; variable boundary conditions.
The stronger terms communicate the graph representation, engineering data structure, prediction target, surrogate purpose, and generalization condition.
Keyword strategy for interdisciplinary articles
Interdisciplinary manuscripts require special care because the same paper must be legible to more than one research community. A purely domain-specific set may hide the methodological novelty, while a purely methodological set may obscure the scientific application.
A useful approach is to allocate keywords across disciplinary interfaces. For an article combining density functional theory, high-throughput computation, and machine learning for battery materials, the set might include:
“density functional theory; materials informatics; high-throughput screening; battery electrode materials; formation energy prediction; uncertainty quantification.”
This set speaks to computational materials science, machine learning, and electrochemical application communities. It also avoids overly broad labels such as “energy” and “artificial intelligence.”
Interdisciplinary terms should be selected based on the journal’s actual readership. A materials journal may require more emphasis on composition, structure, and properties. A machine-learning journal may require the learning problem, representation, and evaluation setting. The same study can justifiably use different keyword sets when submitted to different journals, provided each set remains scientifically accurate.
Result-oriented terms and keyword count
Keywords should identify scientific concepts rather than advertise conclusions. Phrases such as “high accuracy,” “improved performance,” and “superior method” are evaluative and provide little retrieval value. Result-related terms are appropriate only when they name an established phenomenon or outcome, such as “thermal runaway,” “drug resistance,” “fatigue crack growth,” or “phase transition.”
Use the number of keywords required by the journal, and fill optional slots only when each term adds distinct information. A focused methods paper may need four or five terms, whereas a genuinely interdisciplinary article may require six to eight. More terms are not automatically better: an excessive list can dilute semantic focus and introduce peripheral concepts.
The decision can be viewed through marginal information gain:
$
\Delta I_k = I(K \cup \{k\}) - I(K),
$
where $K$ is the current keyword set, $k$ is a candidate term, and $I$ represents the useful information conveyed by the set. When $\Delta I_k$ is negligible because the candidate duplicates an existing term, the candidate should be removed.
Final quality-control checklist
Before submission, verify that the keyword set satisfies the following conditions:
- Every term corresponds to a substantial element of the manuscript.
- The set includes the principal object, problem, method, and context as appropriate.
- Terminology matches current usage in the target research community.
- Controlled vocabulary terms have been checked where relevant.
- Broad, promotional, and ambiguous terms have been removed.
- Abbreviations are standard and unambiguous.
- Near-synonyms do not waste limited keyword slots.
- The set complements rather than mechanically duplicates the title.
- Database searches retrieve literature similar to the manuscript.
- The journal’s exact number and formatting requirements are satisfied.
- The abstract contains the central concepts represented by the keywords.
- The submission-system classifications are consistent with the free-text terms.
If you're working on related challenges in this area and would find guidance helpful, feel free to reach out: CONTACT US.
Conclusion
Choosing keywords for a research article in a Q1 journal is a compact but consequential exercise in scientific classification. The strongest keyword sets are not assembled from generic high-frequency terms or copied from a single related paper. They are derived from the manuscript’s semantic core, aligned with the target journal’s scope, validated against disciplinary vocabulary, and tested in relevant literature databases.
A rigorous process begins by identifying the research object, scientific problem, method, application context, and distinctive contribution. It then expands these concepts into a candidate inventory using recent literature, controlled vocabularies, and accepted terminology. Candidates are evaluated for relevance, specificity, community recognition, journal alignment, retrieval value, and ambiguity. The final set should balance broad field visibility with narrow technical precision while remaining consistent with the title and abstract.
Keywords will not determine whether weak research is accepted by a leading journal. They can, however, influence whether a strong manuscript is classified accurately, routed appropriately, retrieved by the right audience, and integrated into the literature where it belongs. For that reason, keyword selection should be completed with the same precision applied to the title, abstract, figures, and references—not as a last-minute submission formality.
Interested in collaborating on academic research ? feel free to get in touch 🙂.
Check out YouTube channel, published research
you can contact us (bkacademy.in@gmail.com)
Interested to Learn Engineering modelling Check our Courses 🙂