new model · 2026.10.05
AstaBrief 8B: Create a Cited Research Report Locally
Turn retrieved papers into a report whose evidence you can review.
AstaBrief-8B takes a research question and paper excerpts and writes a report with citations. Collect sources with a search tool and pass them to the model to organize arguments and evidence from multiple papers into a draft. The model is released under Apache-2.0 and can run through Transformers or vLLM.
Give it a research question and paper excerpts to get a cited report
Ai2 released `allenai/AstaBrief_8B` under Apache-2.0 on 2026-10-02. The model takes a research question and retrieved scientific-literature excerpts, then writes a report citing the evidence. To compare battery-degradation factors across papers, for example, first use a search tool to collect relevant papers and passages, then send the question and excerpts to AstaBrief. It does not search the web or scholarly databases itself. Ai2 announcement · Model card
First check that the retrieved set contains every comparison target and key condition. Then compare the model’s summaries and citations with the source text so retrieval gaps and writing errors can be fixed separately.
Keep retrieval and report writing as separate steps
Use one task throughout: compare three papers on how high-temperature storage affects lithium-ion battery capacity loss, and tabulate their different test temperatures and observation periods. Keep each retrieved paper’s title, year, DOI or URL, and the actual excerpt. An identifier on each excerpt makes it easier to connect a citation in the report back to its source.
Include the research question, requested report sections, excerpts, and source identifiers. The official SFT prompt has `[QUERY]` and `[SECTION_REFERENCES]` slots, making it straightforward to connect the question with its evidence. AstaBrief model card

Review cited sentences and the scope of their evidence
A citation at the end of a sentence does not mean every part of the sentence appears in the source. If a result from storage at 45°C is generalized to typical degradation at 25°C, the citation may be real while the conclusion has changed scope. Open the cited paper for each comparison and check whether temperature, sample, duration, and measured outcome match.
Ai2 reports citation precision and citation recall as separate measures. Precision concerns whether attached citations support their statements; recall concerns whether necessary evidence for the report’s claims is present. A high score on one does not imply the same score on the other. During review, record unsupported citations and claims with missing citations as different issues.
| Measure | Question | What to do in the battery comparison |
|---|---|---|
| Citation precision | Does an attached citation support the statement? | Check temperature, duration, and measurements in each paper. |
| Citation recall | Has evidence needed for a claim been omitted? | Check that each comparison row has a source identifier. |
Distinguish the two citation measures while reviewing a report sentence.
Citation precision
- Question
- Does an attached citation support the statement?
- What to do in the battery comparison
- Check temperature, duration, and measurements in each paper.
Citation recall
- Question
- Has evidence needed for a claim been omitted?
- What to do in the battery comparison
- Check that each comparison row has a source identifier.

Report quality was evaluated on 100 computer-science questions.
Ai2 evaluated reports on ScholarQA-CS2, a set of 100 user-written computer-science questions. The final AstaBrief-8B averaged 87.0, compared with 77.3 for Qwen3-8B and 83.7 for AstaBrief-8B-SFT in the same table. The evaluation considers coverage of requested content and whether citations support claims. Official model-card evaluation table
Citation precision was 90.5 and citation recall was 78.2. Because recall—which checks whether claims have needed evidence—was lower, begin draft review by finding key claims without citations. These results belong to this question set; evaluate Korean-language sources separately.
| Model | Average score | Evaluation scope |
|---|---|---|
| Qwen3-8B | 77.3 | 100 user-written CS questions |
| AstaBrief-8B-SFT | 83.7 | Same question set |
| AstaBrief-8B | 87.0 | Same question set |
Scores compared in the same ScholarQA-CS2 evaluation table.
Qwen3-8B
- Average score
- 77.3
- Evaluation scope
- 100 user-written CS questions
AstaBrief-8B-SFT
- Average score
- 83.7
- Evaluation scope
- Same question set
AstaBrief-8B
- Average score
- 87.0
- Evaluation scope
- Same question set
Plan memory for weights and the input context.
BF161 typically stores each parameter in two bytes, putting 8B model weights at about 16 GB by simple arithmetic. Longer paper excerpts increase the KV cache2, so check how much memory remains after loading the weights. The official card does not specify a minimum inference3 VRAM4.
First check model loading and response generation using a short excerpt from one paper. Then add papers and excerpts while recording peak memory to find a workable input size for your device.
| Item | Known value | Remaining condition |
|---|---|---|
| Model | 8B, BF16 | No minimum inference VRAM in the card |
| Weight arithmetic | About 16 GB (8B × 2 bytes) | KV cache, runtime, and workspace add memory |
Leave memory for input and generation after loading the weights.
Model
- Known value
- 8B, BF16
- Remaining condition
- No minimum inference VRAM in the card
Weight arithmetic
- Known value
- About 16 GB (8B × 2 bytes)
- Remaining condition
- KV cache, runtime, and workspace add memory
Run the model using the official card’s Python path
The official model card provides Transformers and vLLM Python examples and links to the recommended prompt file. Put the research question and retrieved excerpts in `[QUERY]` and `[SECTION_REFERENCES]`; the code uses the card’s recommended temperature 0.7, top_p 0.95, and maximum 4,096 output tokens5. Batched vLLM input uses `LLM6.generate`. Model card · Official SFT prompt · vLLM generate API
On success, the terminal prints the generated report. If Python cannot find `sft_prompt.txt`, check that you are running from the folder where you downloaded it; for `ModuleNotFoundError`, check that the packages are installed in the active virtual environment7. If loading runs out of memory, do not assume a shorter generation limit solves weight-loading memory; record the GPU8, runtime9, and input context length10, then reassess the precision or model choice. vLLM installation and supported environments
python -m venv .venv
source .venv/bin/activate
python -m pip install -U transformers accelerate vllm
curl -L -o sft_prompt.txt https://huggingface.co/datasets/allenai/AstaBrief_prompts/resolve/main/sft_prompt.txtimport json
from pathlib import Path
from transformers import AutoTokenizer
from vllm import LLM, SamplingParams
model_name = "allenai/AstaBrief_8B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
llm = LLM(model=model_name)
query = "Compare the outcomes reported in these papers and state which conditions differ."
section_references = {
"[P1]": "Paste a relevant quotation from paper 1 here.",
"[P2]": "Paste a relevant quotation from paper 2 here.",
"[P3]": "Paste a relevant quotation from paper 3 here.",
}
formatted_sft_prompt = (
Path("sft_prompt.txt").read_text(encoding="utf-8")
.replace("[QUERY]", query)
.replace("[SECTION_REFERENCES]", json.dumps(section_references, ensure_ascii=False))
)
messages = [{"role": "user", "content": formatted_sft_prompt}]
text = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
sampling = SamplingParams(
temperature=0.7, top_p=0.95, max_tokens=4096,
stop_token_ids=[tokenizer.eos_token_id],
)
outputs = llm.generate([text], sampling)
print(outputs[0].outputs[0].text.strip())
Check the report against all three source papers
For the battery comparison, first check that all three papers appear in the report’s table. Then compare each row’s temperature and observation period with the source, and finally verify that statements in the report point to the papers that support the table values. Keep source identifiers intact so you can trace any mismatch to the original table or passage.
If a paper reports several test temperatures and capacity-measurement times, the model may select just one. Instead of recording only that it was wrong, check whether the excerpts included the condition needed by the research question. If the condition was missing, retrieve the relevant passage; if it was present but changed in the report, correct the cited statement.
| Observed issue | Where to look first | Next action |
|---|---|---|
| A paper needed for comparison is missing | Retrieved results and excerpt list | Revise the query and add relevant passages. |
| Temperature or duration differs from the paper | Cited table and passage | Re-extract the condition and correct the statement. |
| A claim has no citation | Report statement and source identifier | Limit the claim to supported content or add a citation. |
Distinguish retrieval gaps from report-writing errors during review.
A paper needed for comparison is missing
- Where to look first
- Retrieved results and excerpt list
- Next action
- Revise the query and add relevant passages.
Temperature or duration differs from the paper
- Where to look first
- Cited table and passage
- Next action
- Re-extract the condition and correct the statement.
A claim has no citation
- Where to look first
- Report statement and source identifier
- Next action
- Limit the claim to supported content or add a citation.
Move the workflow locally when retrieval and memory are ready
If unpublished research or internal literature must stay off external APIs, you can run open weights on your own hardware. First test retrieval, drafting, and citation review on three public papers; then record model-load memory and end-to-end time with realistic input lengths. Check those results and your organization’s data policy before adding internal material.
For recurring research drafts with human citation review, connect a retrieval system to AstaBrief and test it on three public papers. Use retrieval coverage, conditions in comparison tables, and whether citations support each statement as acceptance checks to identify changes before adopting it in real work.
Terminology notes
BF16 — A 16-bit floating-point format for storing and computing model values. Support depends on the hardware and runtime.
Back to the textKV cache — Memory that stores attention keys and values from earlier tokens for reuse during later token generation. Its size depends on context length and batch size.
Back to the textInference — The process of using a trained model to compute an output for an input. Here, local inference means running the model on the user’s device.
Back to the textVRAM — Memory used by a graphics card’s GPU to store model weights and intermediate values.
Back to the textToken — A unit into which a model divides input or output for processing.
Back to the textLarge language model — A language model trained on large text datasets to process and generate text. Capabilities and supported inputs vary by model.
Back to the textPython virtual environment — An isolated space for installing Python packages per project, helping reduce package-version conflicts.
Back to the textGPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.
Back to the textRuntime — The software environment that provides facilities needed while a program runs. In local AI, it can also refer to a model execution engine.
Back to the textContext window — The token span of input and generated content a model can handle in one request. The supported limit and memory use depend on the model and runtime settings.
Back to the text
Read next
read first
Guide to building a local RAG: Steps to create an AI that answers your documents
Executable programs and extensions
Local LLM·RAG quality evaluation: Verifying changed settings with DeepEval
read first
Local AI for PDFs: check the document before buying more memory
new model
Run Xing4.0 Locally: Install the Coding Model and Plan GPU Memory
