Back
Ai2 open-sources AstaBrief 8B, a fast open-weights model for scientific reports
SiTech AI Team2 min read

Ai2 open-sources AstaBrief 8B, a fast open-weights model for scientific reports

Ai2 has open-sourced AstaBrief 8B, an open-weights model that turns a research question and retrieved literature into a cited report. In Asta's Fast mode it averages 51.1 seconds per report, about 3.5 times faster than the Claude-powered Thinking mode.

Ai2, the Allen Institute for AI, has released AstaBrief 8B, an open-weights model that turns a research question and retrieved literature excerpts into a cited report. It is available in Asta as Fast mode, alongside the Claude-powered Thinking mode.

Across the full Asta pipeline, Fast mode averages 51.1 seconds per report, compared with 178.5 seconds for Thinking mode, about 3.5 times faster. The model writes the full report in one pass, skipping the summarization and clustering stages Thinking mode uses; Ai2 says quality was not sacrificed.

An open model for scientific reports

AstaBrief 8B is an eight-billion-parameter model built on Qwen3-8B. Ai2 focused its effort on post-training data and evaluation rather than pretraining. Open weights also let institutions run the model on their own infrastructure, behind their own firewall.

How it was trained

AstaBrief went through two stages: supervised fine-tuning (SFT) and direct preference optimization (DPO). Ai2 chose this simpler recipe over the costly, unstable reinforcement-learning route it had tried for DR Tulu. Training questions came from real user logs; after filtering, 90K research-focused queries remained.

For SFT, target reports were generated by the ScholarQA pipeline (Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, GPT-4.1), leaving 47K examples after quality filtering. For DPO, report pairs were compared by two judge models (GPT-4.1 and DeepSeek-R1); the final set holds about 6K examples. Of four statistical filters, removing low-citation-density examples delivered the strongest gains.

Results and adoption

The main evaluation target was SQABench-CS2, a set of 200 computer science questions. Ai2 tracked a rubric score, answer precision, citation precision and citation recall, plus DeepScholarBench (63 queries) and a 14-question human study.

AstaBrief results table

AstaBrief was competitive with the Claude-powered pipeline and DR Tulu. In the human study, DR Tulu wins on overall preference, but two of the three researchers preferred AstaBrief on citation accuracy. Ai2 notes the evaluation was mostly done in 2025 and was not rerun against today's models.

LLM-judged comparison chart

Early usage looks encouraging: of 374 Asta users who tried Fast mode, 29.1% kept it for two or more days and users generate an average of 3.67 report threads. Twenty-three percent never switched back; another 18% used both, putting about 40% of their threads through Fast mode. Positive feedback is similar for both modes: 84.2% for Fast and 85.2% for Thinking.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.