new model · 2026.10.05

AstaBrief 8B: Create a Cited Research Report Locally

Turn retrieved papers into a report whose evidence you can review.

AstaBrief-8B takes a research question and paper excerpts and writes a report with citations. Collect sources with a search tool and pass them to the model to organize arguments and evidence from multiple papers into a draft. The model is released under Apache-2.0 and can run through Transformers or vLLM.

Requirements and key details
  • Give the model a research question and retrieved excerpts; a separate tool handles paper search.
  • The official model ID is `allenai/AstaBrief_8B`, under Apache-2.0.
  • On the published 100-question evaluation, citation precision was 90.5 and citation recall was 78.2.
  • The card lists no official minimum memory or local tok/s. Open cited sources and verify that they support the claims.

Give it a research question and paper excerpts to get a cited report

Ai2 released `allenai/AstaBrief_8B` under Apache-2.0 on 2026-10-02. The model takes a research question and retrieved scientific-literature excerpts, then writes a report citing the evidence. To compare battery-degradation factors across papers, for example, first use a search tool to collect relevant papers and passages, then send the question and excerpts to AstaBrief. It does not search the web or scholarly databases itself. Ai2 announcement · Model card

First check that the retrieved set contains every comparison target and key condition. Then compare the model’s summaries and citations with the source text so retrieval gaps and writing errors can be fixed separately.

Keep retrieval and report writing as separate steps

Use one task throughout: compare three papers on how high-temperature storage affects lithium-ion battery capacity loss, and tabulate their different test temperatures and observation periods. Keep each retrieved paper’s title, year, DOI or URL, and the actual excerpt. An identifier on each excerpt makes it easier to connect a citation in the report back to its source.

Include the research question, requested report sections, excerpts, and source identifiers. The official SFT prompt has `[QUERY]` and `[SECTION_REFERENCES]` slots, making it straightforward to connect the question with its evidence. AstaBrief model card

Laptop on a desk with papers and research notes
Retrieval gathers papers and passages; AstaBrief writes from that evidence.

Review cited sentences and the scope of their evidence

A citation at the end of a sentence does not mean every part of the sentence appears in the source. If a result from storage at 45°C is generalized to typical degradation at 25°C, the citation may be real while the conclusion has changed scope. Open the cited paper for each comparison and check whether temperature, sample, duration, and measured outcome match.

Ai2 reports citation precision and citation recall as separate measures. Precision concerns whether attached citations support their statements; recall concerns whether necessary evidence for the report’s claims is present. A high score on one does not imply the same score on the other. During review, record unsupported citations and claims with missing citations as different issues.

Distinguish the two citation measures while reviewing a report sentence.
MeasureQuestionWhat to do in the battery comparison
Citation precisionDoes an attached citation support the statement?Check temperature, duration, and measurements in each paper.
Citation recallHas evidence needed for a claim been omitted?Check that each comparison row has a source identifier.

Distinguish the two citation measures while reviewing a report sentence.

Citation precision

Question
Does an attached citation support the statement?
What to do in the battery comparison
Check temperature, duration, and measurements in each paper.

Citation recall

Question
Has evidence needed for a claim been omitted?
What to do in the battery comparison
Check that each comparison row has a source identifier.
Three sets of papers and writing tools spread across a desk
Open the citation and check that conditions and claim scope match the source.

Report quality was evaluated on 100 computer-science questions.

Ai2 evaluated reports on ScholarQA-CS2, a set of 100 user-written computer-science questions. The final AstaBrief-8B averaged 87.0, compared with 77.3 for Qwen3-8B and 83.7 for AstaBrief-8B-SFT in the same table. The evaluation considers coverage of requested content and whether citations support claims. Official model-card evaluation table

Citation precision was 90.5 and citation recall was 78.2. Because recall—which checks whether claims have needed evidence—was lower, begin draft review by finding key claims without citations. These results belong to this question set; evaluate Korean-language sources separately.

Scores compared in the same ScholarQA-CS2 evaluation table.
ModelAverage scoreEvaluation scope
Qwen3-8B77.3100 user-written CS questions
AstaBrief-8B-SFT83.7Same question set
AstaBrief-8B87.0Same question set

Scores compared in the same ScholarQA-CS2 evaluation table.

Qwen3-8B

Average score
77.3
Evaluation scope
100 user-written CS questions

AstaBrief-8B-SFT

Average score
83.7
Evaluation scope
Same question set

AstaBrief-8B

Average score
87.0
Evaluation scope
Same question set

Plan memory for weights and the input context.

BF161 typically stores each parameter in two bytes, putting 8B model weights at about 16 GB by simple arithmetic. Longer paper excerpts increase the KV cache2, so check how much memory remains after loading the weights. The official card does not specify a minimum inference3 VRAM4.

First check model loading and response generation using a short excerpt from one paper. Then add papers and excerpts while recording peak memory to find a workable input size for your device.

Leave memory for input and generation after loading the weights.
ItemKnown valueRemaining condition
Model8B, BF16No minimum inference VRAM in the card
Weight arithmeticAbout 16 GB (8B × 2 bytes)KV cache, runtime, and workspace add memory

Leave memory for input and generation after loading the weights.

Model

Known value
8B, BF16
Remaining condition
No minimum inference VRAM in the card

Weight arithmetic

Known value
About 16 GB (8B × 2 bytes)
Remaining condition
KV cache, runtime, and workspace add memory

Run the model using the official card’s Python path

The official model card provides Transformers and vLLM Python examples and links to the recommended prompt file. Put the research question and retrieved excerpts in `[QUERY]` and `[SECTION_REFERENCES]`; the code uses the card’s recommended temperature 0.7, top_p 0.95, and maximum 4,096 output tokens5. Batched vLLM input uses `LLM6.generate`. Model card · Official SFT prompt · vLLM generate API

On success, the terminal prints the generated report. If Python cannot find `sft_prompt.txt`, check that you are running from the folder where you downloaded it; for `ModuleNotFoundError`, check that the packages are installed in the active virtual environment7. If loading runs out of memory, do not assume a shorter generation limit solves weight-loading memory; record the GPU8, runtime9, and input context length10, then reassess the precision or model choice. vLLM installation and supported environments

Prepare the runtime and official prompt file
python -m venv .venv
source .venv/bin/activate
python -m pip install -U transformers accelerate vllm
curl -L -o sft_prompt.txt https://huggingface.co/datasets/allenai/AstaBrief_prompts/resolve/main/sft_prompt.txt
Install vLLM in a Python environment with a supported GPU and driver. Before installing, check the [official vLLM installation guide](https://docs.vllm.ai/en/stable/getting_started/installation/) and model card.
Generate one report from retrieved excerpts
import json
from pathlib import Path

from transformers import AutoTokenizer
from vllm import LLM, SamplingParams

model_name = "allenai/AstaBrief_8B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
llm = LLM(model=model_name)

query = "Compare the outcomes reported in these papers and state which conditions differ."
section_references = {
    "[P1]": "Paste a relevant quotation from paper 1 here.",
    "[P2]": "Paste a relevant quotation from paper 2 here.",
    "[P3]": "Paste a relevant quotation from paper 3 here.",
}
formatted_sft_prompt = (
    Path("sft_prompt.txt").read_text(encoding="utf-8")
    .replace("[QUERY]", query)
    .replace("[SECTION_REFERENCES]", json.dumps(section_references, ensure_ascii=False))
)
messages = [{"role": "user", "content": formatted_sft_prompt}]
text = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)
sampling = SamplingParams(
    temperature=0.7, top_p=0.95, max_tokens=4096,
    stop_token_ids=[tokenizer.eos_token_id],
)
outputs = llm.generate([text], sampling)
print(outputs[0].outputs[0].text.strip())
Put the actual question and literature excerpts into the two slots in `sft_prompt.txt`. `[P1]` and `[P2]` are example source IDs. The snippet sends instructional placeholders if run unchanged, so replace them with the material for your report.
Report open on a laptop beside papers and comparison notes
Review each comparison value alongside the paper that supports it.

Check the report against all three source papers

For the battery comparison, first check that all three papers appear in the report’s table. Then compare each row’s temperature and observation period with the source, and finally verify that statements in the report point to the papers that support the table values. Keep source identifiers intact so you can trace any mismatch to the original table or passage.

If a paper reports several test temperatures and capacity-measurement times, the model may select just one. Instead of recording only that it was wrong, check whether the excerpts included the condition needed by the research question. If the condition was missing, retrieve the relevant passage; if it was present but changed in the report, correct the cited statement.

Distinguish retrieval gaps from report-writing errors during review.
Observed issueWhere to look firstNext action
A paper needed for comparison is missingRetrieved results and excerpt listRevise the query and add relevant passages.
Temperature or duration differs from the paperCited table and passageRe-extract the condition and correct the statement.
A claim has no citationReport statement and source identifierLimit the claim to supported content or add a citation.

Distinguish retrieval gaps from report-writing errors during review.

A paper needed for comparison is missing

Where to look first
Retrieved results and excerpt list
Next action
Revise the query and add relevant passages.

Temperature or duration differs from the paper

Where to look first
Cited table and passage
Next action
Re-extract the condition and correct the statement.

A claim has no citation

Where to look first
Report statement and source identifier
Next action
Limit the claim to supported content or add a citation.

Move the workflow locally when retrieval and memory are ready

If unpublished research or internal literature must stay off external APIs, you can run open weights on your own hardware. First test retrieval, drafting, and citation review on three public papers; then record model-load memory and end-to-end time with realistic input lengths. Check those results and your organization’s data policy before adding internal material.

For recurring research drafts with human citation review, connect a retrieval system to AstaBrief and test it on three public papers. Use retrieval coverage, conditions in comparison tables, and whether citations support each statement as acceptance checks to identify changes before adopting it in real work.

Terminology notes

  1. BF16 — A 16-bit floating-point format for storing and computing model values. Support depends on the hardware and runtime.

    Back to the text
  2. KV cache — Memory that stores attention keys and values from earlier tokens for reuse during later token generation. Its size depends on context length and batch size.

    Back to the text
  3. Inference — The process of using a trained model to compute an output for an input. Here, local inference means running the model on the user’s device.

    Back to the text
  4. VRAM — Memory used by a graphics card’s GPU to store model weights and intermediate values.

    Back to the text
  5. Token — A unit into which a model divides input or output for processing.

    Back to the text
  6. Large language model — A language model trained on large text datasets to process and generate text. Capabilities and supported inputs vary by model.

    Back to the text
  7. Python virtual environment — An isolated space for installing Python packages per project, helping reduce package-version conflicts.

    Back to the text
  8. GPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.

    Back to the text
  9. Runtime — The software environment that provides facilities needed while a program runs. In local AI, it can also refer to a model execution engine.

    Back to the text
  10. Context window — The token span of input and generated content a model can handle in one request. The supported limit and memory use depend on the model and runtime settings.

    Back to the text