What happened?
OpenAI introduced GeneBench-Pro to evaluate the performance of AI systems in long-term genome analysis and quantitative biology tasks.
A standard has been established to evaluate the performance of scientific AI not through short-answer questions, but through long-term tasks closer to actual research workflows. The content of the announcement was organized based on official sources, and in actual use, the scope of provision and technical limitations should be reviewed together.
Why does it matter?
A standard has been established to evaluate the performance of scientific AI not through short-answer questions, but through long-term tasks closer to actual research workflows.