nit: standardize naming (#849)

use correct project name in a few places that were missed
This commit is contained in:
Mason Daugherty
2026-01-20 11:53:42 -05:00
committed by GitHub
parent 319f105ff1
commit 16f31d09a6
5 changed files with 35 additions and 19 deletions
+1 -1
View File
@@ -16,7 +16,7 @@ __pycache__/
*.so
.Python
# DeepAgent filesystem (if using local backend)
# Deep Agent filesystem (if using local backend)
.deepagent_fs/
agent_files/
agent_workspace/
+25 -11
View File
@@ -26,18 +26,21 @@ Uses the [Chinook database](https://github.com/lerocha/chinook-database) - a sam
### Installation
1. Clone the deepagents repository and navigate to this example:
```bash
git clone https://github.com/langchain-ai/deepagents.git
cd deepagents/examples/text-to-sql-agent
```
2. Download the Chinook database:
1. Download the Chinook database:
```bash
# Download the SQLite database file
curl -L -o chinook.db https://github.com/lerocha/chinook-database/raw/master/ChinookDatabase/DataSources/Chinook_Sqlite.sqlite
```
3. Create a virtual environment and install dependencies:
1. Create a virtual environment and install dependencies:
```bash
# Using uv (recommended)
uv venv --python 3.11
@@ -45,18 +48,21 @@ source .venv/bin/activate # On Windows: .venv\Scripts\activate
uv pip install -e .
```
4. Set up your environment variables:
1. Set up your environment variables:
```bash
cp .env.example .env
# Edit .env and add your API keys
```
Required in `.env`:
```
ANTHROPIC_API_KEY=your_anthropic_api_key_here
```
Optional:
```
LANGCHAIN_TRACING_V2=true
LANGSMITH_ENDPOINT=https://api.smith.langchain.com
@@ -100,15 +106,14 @@ result = agent.invoke({
print(result["messages"][-1].content)
```
## How DeepAgent Works
## How the Deep Agent Works
### Architecture
```
User Question
↓
DeepAgent (with planning)
Deep Agent (with planning)
├─ write_todos (plan the approach)
├─ SQL Tools
│ ├─ list_tables
@@ -129,15 +134,17 @@ Formatted Answer
### Configuration
DeepAgent uses **progressive disclosure** with memory files and skills:
Deep Agents uses **progressive disclosure** with memory files and skills:
**AGENTS.md** (always loaded) - Contains:
- Agent identity and role
- Core principles and safety rules
- General guidelines
- Communication style
**skills/** (loaded on-demand) - Specialized workflows:
- **query-writing** - How to write and execute SQL queries (simple and complex)
- **schema-exploration** - How to discover database structure and relationships
@@ -146,25 +153,30 @@ The agent sees skill descriptions in its context but only loads the full SKILL.m
## Example Queries
### Simple Query
```
"How many customers are from Canada?"
```
The agent will directly query and return the count.
### Complex Query with Planning
```
"Which employee generated the most revenue and from which countries?"
```
The agent will:
1. Use `write_todos` to plan the approach
2. Identify required tables (Employee, Invoice, Customer)
3. Plan the JOIN structure
4. Execute the query
5. Format results with analysis
## DeepAgent Output Example
## Deep Agent Output Example
DeepAgent shows its reasoning process:
The Deep Agent shows its reasoning process:
```
Question: Which employee generated the most revenue by country?
@@ -230,6 +242,7 @@ All dependencies are specified in `pyproject.toml`:
1. Sign up for a free account at [LangSmith](https://smith.langchain.com/)
2. Create an API key from your account settings
3. Add these variables to your `.env` file:
```
LANGCHAIN_TRACING_V2=true
LANGSMITH_ENDPOINT=https://api.smith.langchain.com
@@ -241,9 +254,10 @@ LANGCHAIN_PROJECT=text2sql-deepagent
When configured, every query is automatically traced:
![DeepAgent LangSmith Trace Example](text-to-sql-langsmith-trace.png)
![Deep Agent LangSmith Trace Example](text-to-sql-langsmith-trace.png)
You can view:
- Complete execution trace with all tool calls
- Planning steps (write_todos)
- Filesystem operations
@@ -251,7 +265,7 @@ You can view:
- Generated SQL queries
- Error messages and retry attempts
View your traces at: https://smith.langchain.com/
View your traces at: <https://smith.langchain.com/>
## Resources
@@ -1,4 +1,4 @@
"""Middleware for the DeepAgent."""
"""Middleware for the agent."""
from deepagents.middleware.filesystem import FilesystemMiddleware
from deepagents.middleware.memory import MemoryMiddleware
+7 -5
View File
@@ -1,8 +1,8 @@
# Building DeepAgent Harnesses for Terminal Bench 2.0 with Harbor
# Building Deep Agent Harnesses for Terminal Bench 2.0 with Harbor
## Overview
This repository demonstrates how to evaluate and improve your DeepAgent harness using [Harbor](https://github.com/laude-institute/harbor) and [LangSmith](https://smith.langchain.com).
This repository demonstrates how to evaluate and improve your Deep Agent harness using [Harbor](https://harborframework.com/) and [LangSmith](https://www.langchain.com/langsmith/observability).
### What is Harbor?
@@ -11,21 +11,22 @@ Harbor is an evaluation framework that simplifies running agents on challenging
- **Sandbox environments** (Docker, Modal, Daytona, E2B, etc.)
- **Automatic test execution** and verification
- **Reward scoring** (0.0 - 1.0 based on test pass rate)
- **Trajectory logging** in ATIF format (Agent Trajectory Interchange Format)
- **Trajectory logging** in ATIF format [(Agent Trajectory Interchange Format)](https://harborframework.com/docs/trajectory-format)
### What is Terminal Bench 2.0?
[Terminal Bench 2.0](https://github.com/laude-institute/terminal-bench-2) is an evaluation benchmark that measures agent capabilities across several domains, testing how well an agent operates using a computer environment, primarily via the terminal. The benchmark includes 90+ tasks across domains like software engineering, biology, security, gaming, and more.
**Example tasks:**
- `path-tracing`: Reverse-engineer C program from rendered image
- `chess-best-move`: Find optimal move using chess engine
- `git-multibranch`: Complex git operations with merge conflicts
- `sqlite-with-gcov`: Build SQLite with code coverage, analyze reports
### The DeepAgent Architecture
### The Deep Agent Architecture
The DeepAgent harness ships with design patterns validated as good defaults across agentic tasks:
The Deep Agent harness ships with design patterns validated as good defaults across agentic tasks:
1. **Detailed System Prompt**: Expansive, instructional prompts with tool guidance and examples
2. **Planning Middleware**: The `write_todos` tool helps the agent structure thinking and track progress
@@ -148,6 +149,7 @@ Harbor supports multiple sandbox environments. Use the `--env` flag to select:
- `runloop` - Runloop sandboxes
Makefile shortcuts are available for common workflows:
- `make run-terminal-bench-docker` - Run 1 task locally with Docker
- `make run-terminal-bench-daytona` - Run 10 tasks on Daytona
- `make run-terminal-bench-modal` - Run 4 tasks on Modal
@@ -156,7 +156,7 @@ class DeepAgentsWrapper(BaseAgent):
environment: BaseEnvironment,
context: AgentContext,
) -> None:
"""Execute the DeepAgent on the given instruction.
"""Execute the Deep Agent on the given instruction.
Args:
instruction: The task to complete