Capabilities of Large Language Models in Physics Research Planning and Proposal Evaluation: A Human-AI Blind Test Comparison

This study investigates the practical capabilities of large language models (LLMs) in assisting with physics research planning and proposal evaluation. Eight expert-level research projects from the fields of physics, astrophysics, and cosmology were selected. For each project, a one-page research proposal was independently generated by human research teams and three mainstream LLMs—ChatGPT, Claude, and DeepSeek—yielding 32 proposals in total. These proposals were then anonymously evaluated by four human reviewers and two frontier LLMs (Claude Opus 4.8 and ChatGPT Pro 5.5) across four dimensions: score quality, research significance, feasibility, and innovation, and the evaluators were also asked to identify whether each proposal was written by a human or an AI. Results show that human reviewers gave overall comparable scores to both AI-generated and human-authored proposals. However, AI reviewers awarded significantly higher scores to AI-generated proposals, approximately one point higher on average. In the authorship identification task, human reviewers achieved 72% and 79% accuracy for AI and human authors respectively, while AI reviewers reached 100% accuracy. These findings reveal both the potential and the systematic bias inherent in current LLMs when used for scientific research assistance, suggesting that caution should be exercised as LLMs are increasingly deployed in research workflows.

Background and Context

The integration of Large Language Models (LLMs) into the scientific workflow has evolved from simple literature summarization to complex research assistance tasks. This study addresses a critical gap in understanding whether LLMs can effectively assist in physics research planning and proposal evaluation. The research focuses on the fields of physics, astrophysics, and cosmology, selecting eight expert-level research projects as the foundation for the experiment. The primary objective was to simulate a real-world scientific grant application scenario to evaluate the quality of AI-generated proposals compared to human-authored ones.

To achieve this, the research team employed a rigorous blind-test methodology. For each of the eight selected projects, one-page research proposals were independently generated by human research teams and three mainstream LLMs: ChatGPT, Claude, and DeepSeek. These models represent the state-of-the-art as of mid-2025, with tool access enabled by default to ensure realistic performance conditions. This process yielded a total of 32 proposals, creating a diverse dataset that allowed for a direct comparison between human expertise and AI-generated content in a constrained format.

The evaluation phase involved both human reviewers and advanced AI models. Four human reviewers and two frontier LLMs, specifically Claude Opus 4.8 and ChatGPT Pro 5.5, assessed the proposals anonymously. The assessment was conducted across four key dimensions: score quality, research significance, feasibility, and innovation. Additionally, all evaluators were tasked with identifying whether each proposal was written by a human or an AI. This dual approach not only measured the quality of the proposals but also tested the perceptual capabilities of both humans and AI in detecting machine-generated text, providing a comprehensive view of the current capabilities and limitations of LLMs in scientific contexts.

Deep Analysis

The results of the evaluation reveal significant disparities in how human and AI reviewers assess content. Human reviewers awarded comparable scores to both AI-generated and human-authored proposals, indicating that current LLMs can produce proposals that are professionally complete and logically consistent. In the context of one-page summaries, AI output is now competitive with human experts in terms of structure and adherence to professional norms. However, when AI models were used as reviewers, a distinct bias emerged. The two frontier LLMs awarded significantly higher scores to AI-generated proposals, averaging approximately one point higher on a five-point scale compared to human-authored ones.

This "in-group preference" among AI reviewers highlights a systemic bias that could distort scientific evaluation processes if left unchecked. The tendency of AI to favor its own kind suggests that automated review systems might inadvertently create an echo chamber, where AI-generated content is disproportionately valued over human insights. This bias is not merely a minor statistical fluctuation but a structural characteristic of current LLMs that could impact the allocation of research resources and the recognition of scientific contributions.

Furthermore, the authorship identification task yielded striking results that underscore the sophistication of modern LLMs. While human reviewers achieved accuracy rates of 72% and 79% in identifying AI and human authors respectively, the AI reviewers reached a perfect 100% accuracy rate. This indicates that AI models have developed an acute sensitivity to the stylistic and structural markers of machine-generated text. The ability of AI to perfectly distinguish between human and AI writing suggests that the boundary between human and machine authorship is becoming increasingly discernible to algorithms, even if it remains challenging for human reviewers. This finding has profound implications for academic integrity and the verification of scientific authorship.

Industry Impact

The findings of this study have immediate implications for the deployment of LLMs in scientific research workflows. On one hand, the ability of LLMs to generate high-quality, structured proposals positions them as valuable collaborators in the early stages of research planning. They can serve as efficient tools for brainstorming, outlining, and refining ideas, potentially accelerating the initial phases of grant applications and project planning. This capability is particularly relevant for researchers who may struggle with the administrative aspects of scientific communication, allowing them to focus more on the core scientific questions.

On the other hand, the identified bias in AI reviewers poses a significant risk to the fairness and integrity of scientific evaluation. If LLMs are integrated into peer review or grant assessment processes without proper safeguards, the systematic preference for AI-generated content could lead to a homogenization of scientific discourse. This could marginalize human creativity and innovative approaches that do not conform to the patterns learned by AI models. The potential for an "echo chamber" effect necessitates the development of robust mechanisms to detect and mitigate such biases, ensuring that evaluation processes remain fair and objective.

For the open-source community and AI developers, this study provides a critical benchmark for evaluating the fairness and robustness of LLMs in specialized domains. It highlights the need for developers to prioritize the reduction of systemic biases in model training and evaluation. Future iterations of LLMs should be tested not only for their generative capabilities but also for their ability to provide unbiased assessments. This requires a shift in development priorities towards ensuring that AI tools enhance rather than distort the scientific process.

Outlook

Looking ahead, the integration of LLMs into scientific research will require a nuanced approach that balances efficiency with integrity. One promising direction is the development of hybrid review mechanisms that combine the speed and consistency of AI with the nuanced judgment of human reviewers. By leveraging the strengths of both, scientific communities can mitigate the risks of AI bias while still benefiting from the productivity gains offered by automation. Additionally, further research is needed to explore how prompt engineering and fine-tuning techniques can be used to reduce the in-group preference observed in AI reviewers.

Another area of future investigation is the development of more sophisticated evaluation metrics that go beyond simple scoring. Current metrics may not fully capture the depth and originality of scientific contributions, especially when comparing human and AI-generated content. New frameworks that can assess the novelty, impact, and methodological rigor of proposals in a more holistic manner will be essential for fair evaluation. Moreover, as AI models continue to evolve, ongoing monitoring and adaptation of evaluation protocols will be necessary to address emerging biases and challenges.

Ultimately, the role of LLMs in scientific research should be viewed as complementary rather than substitutive. While AI can enhance the efficiency of research planning and proposal writing, the core values of scientific inquiry—curiosity, creativity, and critical thinking—remain distinctly human. By maintaining a clear distinction between the supportive role of AI and the authoritative role of human judgment, the scientific community can harness the power of LLMs while preserving the integrity and diversity of scientific discourse. This balanced approach will be crucial for ensuring that AI serves as a tool for empowerment rather than a source of distortion in the pursuit of knowledge.

Sources