Use your own compute
Start independent research with CiteArk Agent and optionally upload signed CAP artifacts.
CiteArk Agent reads papers and code, plans experiments for your goals and available compute, executes them, evaluates evidence and signs CAP artifacts. You can use it independently. CiteArk also hosts the same engine and stores research artifacts.
Start
Install Node.js 22 or newer, Git and tar. Local execution requires Docker. Windows uses Docker Desktop with WSL2 and Linux containers; Apple Silicon supports CPU execution. CUDA requires compatible NVIDIA hardware. You can also run on a Linux server over SSH, or supply a GCP Batch or AutoDL catalog with --compute-catalog.
Install once
Copy this command into your terminal to install or update CiteArk Agent:
npm install -g https://citeark.co/cli.tgzAfter installation, start with:
citearkThe interactive entry accepts a paper file, paper URL, CAP or an existing workspace. It asks for your goal, compute catalog and initial model configuration, then presents experiments for selection. Saved model settings are reused.
On a paper page, choose Use your compute and run the second command, citeark start <paper-reference>, to start from that paper's materials. A previous successful platform run is not required.
If PowerShell blocks scripts, use npm.cmd and citeark.cmd. If citeark is not found after installation, reopen your terminal and check that npm's global command directory is on PATH.
You can also start without CiteArk:
citeark start --paper ./paper.pdf --repository ./code --goal 'Compare inference batch sizes'
citeark start --cap ./previous.cap --paper ./paper.pdf --guide ./REPRODUCE.mdA CAP can provide authenticated source understanding, historical evidence and reference code. The Agent still plans for this study. Instructions help find a route; they do not fix the experiments or replace fresh observations. Explicit replay of an existing delivery remains available as deploy <paper-reference> --replay.
Use --plan-only to save a plan without executing experiments. Resume with start --resume --work-dir <directory> --experiments <id1,id2>. A new goal or input starts a new workspace. --dry-run prepares inputs without model calls, container builds or experiments.
Parse PDFs on your own machine
Before reading a PDF, install Python 3.12 and unzip, then run:
citeark paper setup --python python3.12
citeark paper parse --input ./paper.pdf --output ./paper-readingSetup installs a pinned parser in an isolated environment and downloads public models. The second command tests parsing without calling a research model or starting experiments. Normal citeark start uses the same parser. It defaults to four CPU threads; 16 GiB of memory is recommended. No GPU, CiteArk account, Google Cloud project, or MinerU API key is required. Use WSL2 for PDF parsing on Windows; native Windows parsing has not been verified.
Models are stored in ~/.citeark/mineru-4.0.10 and reused offline. Set CITEARK_MINERU_HOME to change the directory, or add --source modelscope during setup to download from ModelScope. Linux also needs libgl1, libglib2.0-0, libgomp1, and libvulkan1. For SSH execution, complete setup on the server.
The reading copy retains physical page locations, equations, tables, and images. Digest-based caching avoids repeated parsing. Keep the original PDF: subscripts and caption associations can still be wrong and need page-level inspection. Hosted on-demand CPU tasks execute the same parser shipped in the Agent package; cloud transport, queues, and spending belong to the host.
Follow the research session
The English terminal interface shows five stages: Get paper, Understand research, Prepare materials, Run experiments, and Assess & deliver. The conversation displays the Agent's actual user-facing messages, tool calls, commands and results as they arrive. Assistant messages remain complete; Ctrl+O expands long tool results.
Use the mouse wheel or PgUp/PgDn to browse history. New output does not move your view while you are reading earlier messages; End returns to live output. Runtime events determine when text becomes visible, so some tools emit a result only after finishing.
The workspace saves session.events.jsonl. Open it without running research:
citeark view ./my-researchInteractive research runs in a separate worker. Ctrl+C detaches the terminal; reconnect with citeark session attach ./my-research. Enter follow-up instructions while attached; they apply at the next Agent invocation. /pause pauses at the next stage boundary, /resume continues a paused worker, and /cancel requests cancellation. Inspect citeark session status ./my-research for status and resource cleanup results.
Use --plain for ordinary output. Add --background to a noninteractive start to keep it running after disconnection. citeark start --resume --work-dir ./my-research restarts an interrupted workspace; completed research does not rerun. Training recovery depends on its saved checkpoints.
Models and servers
Save reusable model and compute profiles:
citeark model add research --agent opencode --api-base-url https://openrouter.ai/api/v1 --model <model-id> --api-key-env RESEARCH_API_KEY
citeark model use research
citeark compute add lab --type ssh --host my-lab --directory /home/research/citeark
citeark start --paper ./paper.pdf --compute labSet the model key in the named environment variable. model check research tests tool calling and structured output with up to three small, billable requests. The interactive /model and /compute commands select saved profiles. Model and compute snapshots belong to each workspace; changing defaults does not alter existing research. --assessment-profile can choose a separate scientific assessor.
The SSH server needs Node.js, CiteArk Agent, a configured model, and Docker for local execution. Configure its credentials on the server. Use an existing SSH host alias and verified host key; compute check lab checks the connection. The remote worker returns its workspace path. Reconnect with session attach <remote-workspace> --compute lab, choose experiments with start --resume --work-dir <remote-workspace> --compute lab --experiments <ids>, and retrieve verified CAP files with session fetch <remote-workspace> --compute lab --output ./results.
User servers and CiteArk servers run the same Agent. The platform adds account, queue and billing services. The project remains alpha: protocol and process tests do not establish real-paper or GPU reliability. --model-budget is a runtime limit, not a total study cost cap.
Install from source archive
Download the Agent archive, extract it into a new directory and run:
npm ci
node src/cli.mjs setup
node src/cli.mjs paper setup --python python3.12
node src/cli.mjs start --paper ./paper.pdfsetup --gpu builds the CUDA runtime. The website pins download archives to an Agent commit and digest. The current archives contain JavaScript source and require Node.js. The stable /cli.tgz address follows the current website release. Run the install command again to upgrade; an installed Agent does not update itself.
Results and upload
The workspace retains planning, execution evidence and checkpoints. Completed CAPs are listed in the terminal and in workspace.json. The default signing key is local at ~/.citeark/signing-key.pem; --signing-key selects another Ed25519 key.
Local use requires no platform account. Optional upload uses a platform API key with write permission:
citeark publish --plan ./plan.cap --cap ./result.cap
citeark publish --plan ./plan.cap --cap ./result.cap --repository-id '<paper-id>' --visibility publicUpload the signed plan before its execution result. Independent artifacts need no existing platform run. Without a repository, uploads default to private; --visibility public explicitly publishes them. Community results keep their own experiments, uploader and signing key and do not overwrite hosted assessments.
GET /api/v1/artifacts lists your uploads. GET /api/v1/artifacts/{artifactDigest} returns metadata; append ?download to retrieve the CAP. POST /api/v1/artifacts?visibility=private accepts the binary CAP body, with optional repository=<id> and explicit visibility=public. Existing paper community upload endpoints remain available.
Uploads are limited to 24 MiB compressed and 96 MiB expanded. Dataset and model bytes must stay at their original sources or on your compute. A valid signature establishes provenance of the artifact, while its assessment describes what the evidence supports.
Inspect research routes
CiteArk Agent produces CAP 2.0.0-alpha.2. citeark graph --cap result.cap --operation subgraph verifies and reads a graph offline; use --query query.json for exact route, comparison or provenance selections. The repository Research routes tab and account menu Research artifacts page use the same model. Paper-free graphs can be uploaded and default to private. A full native research editor is deferred.