Compare commits
55
Commits
@@ -15,6 +15,7 @@ docs/
|
||||
*.pyzz
|
||||
.venv/
|
||||
venv/
|
||||
.worktrees/
|
||||
__pycache__/
|
||||
poetry.lock
|
||||
.pytest_cache/
|
||||
|
||||
@@ -0,0 +1,265 @@
|
||||
# Design: Native Anthropic Tools Integration
|
||||
|
||||
**Goal**: Integrate Anthropic's native trained tools (bash_20250124, text_editor_20250728, computer_20251124) into nanobot to leverage model's trained behaviors instead of custom function tools.
|
||||
|
||||
## Overview
|
||||
|
||||
Anthropic's native tools are version-coupled to model training. Unlike custom function tools (which the model learns via instruction-following at inference time), native tools have their behaviors baked into model weights during training. This provides more reliable tool execution.
|
||||
|
||||
**Key Insight**: The Anthropic API accepts BOTH tool formats in the same request:
|
||||
- Function tools: `{type: "function", function: {name, description, input_schema}}`
|
||||
- Native tools: `{type: "bash_20250124", name: "bash"}` (schema-less)
|
||||
|
||||
## Architecture
|
||||
|
||||
### 1. Tool Addition Strategy
|
||||
|
||||
Add three native tool implementations from anthropic-quickstarts reference:
|
||||
- **BashTool20250124** - persistent bash session (replaces ExecTool)
|
||||
- **EditTool20250728** - file operations with view/create/str_replace/insert (replaces EditTool, possibly ReadFileTool/WriteFileTool)
|
||||
- **ComputerTool20251124** - VNC desktop control (new capability)
|
||||
|
||||
Location: `nanobot/agent/tools/anthropic/` (new subpackage)
|
||||
|
||||
Port from reference:
|
||||
- Base classes: `BaseAnthropicTool`, `ToolResult`, `CLIResult`, `ToolError`
|
||||
- Tool implementations with trained behaviors intact
|
||||
- Session management (_BashSession for bash tool)
|
||||
|
||||
### 2. Registry Changes
|
||||
|
||||
Make `ToolRegistry` format-agnostic via duck typing:
|
||||
|
||||
**Current**: Only calls `tool.to_schema()`, expects function format
|
||||
|
||||
**New**: Support both interfaces
|
||||
```python
|
||||
def get_definitions(self) -> list[dict[str, Any]]:
|
||||
definitions = []
|
||||
for tool in self._tools.values():
|
||||
if hasattr(tool, 'to_params'): # Native Anthropic tool
|
||||
definitions.append(tool.to_params())
|
||||
elif hasattr(tool, 'to_schema'): # Function tool
|
||||
definitions.append(tool.to_schema())
|
||||
else:
|
||||
raise ValueError(f"Tool {tool.name} has no schema method")
|
||||
return definitions
|
||||
```
|
||||
|
||||
**Execution**: No changes needed - `execute()` already looks up by name and calls the tool. Native tools implement `__call__(**kwargs)` which works with existing dispatch.
|
||||
|
||||
**Result**: Registry becomes thin coordination layer, doesn't enforce specific base class.
|
||||
|
||||
### 3. Tool Implementations
|
||||
|
||||
#### BashTool20250124
|
||||
- Maintains persistent bash session via `_BashSession` class
|
||||
- Sentinel-based output reading for reliable command capture
|
||||
- Timeout handling (120s default)
|
||||
- Restart capability
|
||||
- Returns: `ToolResult(output=..., error=...)`
|
||||
|
||||
#### EditTool20250728
|
||||
- Commands: `view`, `create`, `str_replace`, `insert`
|
||||
- Path validation (absolute paths required)
|
||||
- `str_replace`: uniqueness checking before replacement
|
||||
- `insert`: line number validation
|
||||
- File history tracking for potential undo
|
||||
- Returns: `CLIResult(output=...)` with formatted snippets
|
||||
|
||||
#### ComputerTool20251124
|
||||
- VNC desktop interaction (keyboard, mouse, screenshots)
|
||||
- Actions: `key`, `type`, `mouse_move`, `left_click`, `right_click`, `double_click`, `screenshot`, etc.
|
||||
- Screenshot returns `ToolResult(base64_image=...)`
|
||||
- Coordinate scaling support
|
||||
- Connects to VNC at 172.17.0.1:5900 (Windows VM from code-server)
|
||||
|
||||
### 4. API Integration
|
||||
|
||||
Update `anthropic_oauth.py._convert_tools_to_anthropic()` to pass through both formats:
|
||||
|
||||
**Current**: Only converts `type: "function"` tools
|
||||
```python
|
||||
if tool.get("type") == "function":
|
||||
# convert to Anthropic format
|
||||
```
|
||||
|
||||
**New**: Pass through ALL formats
|
||||
```python
|
||||
def _convert_tools_to_anthropic(self, tools: list[dict[str, Any]] | None) -> list[dict[str, Any]] | None:
|
||||
if not tools:
|
||||
return None
|
||||
|
||||
anthropic_tools = []
|
||||
for tool in tools:
|
||||
if tool.get("type") == "function":
|
||||
# Convert function tool format
|
||||
func = tool["function"]
|
||||
anthropic_tools.append({
|
||||
"name": func["name"],
|
||||
"description": func.get("description", ""),
|
||||
"input_schema": func.get("parameters", {"type": "object", "properties": {}})
|
||||
})
|
||||
else:
|
||||
# Pass through native tool format as-is
|
||||
# (bash_20250124, text_editor_20250728, computer_20251124)
|
||||
anthropic_tools.append(tool)
|
||||
|
||||
return anthropic_tools if anthropic_tools else None
|
||||
```
|
||||
|
||||
**Distinction**: Based on `type` field
|
||||
- `type == "function"` → function tool, needs conversion
|
||||
- `type == "bash_20250124"` (or other native type) → pass through as-is
|
||||
|
||||
### 5. Tool Result Handling
|
||||
|
||||
**Current**: Tools return plain strings
|
||||
|
||||
**New**: Native tools return `ToolResult` objects
|
||||
```python
|
||||
@dataclass(kw_only=True, frozen=True)
|
||||
class ToolResult:
|
||||
output: str | None = None
|
||||
error: str | None = None
|
||||
base64_image: str | None = None
|
||||
system: str | None = None
|
||||
```
|
||||
|
||||
**Agent loop changes** (`loop.py`): Handle both return types
|
||||
```python
|
||||
result = await self.tools.execute(tool_name, tool_input)
|
||||
|
||||
if isinstance(result, ToolResult):
|
||||
# Native tool result - build structured content
|
||||
tool_result_content = []
|
||||
if result.output:
|
||||
tool_result_content.append({"type": "text", "text": result.output})
|
||||
if result.error:
|
||||
tool_result_content.append({"type": "text", "text": f"Error: {result.error}"})
|
||||
if result.base64_image:
|
||||
# Image handling (see Section 6)
|
||||
pass
|
||||
if result.system:
|
||||
# System messages for next turn
|
||||
pass
|
||||
else:
|
||||
# Legacy string result from function tools
|
||||
tool_result_content = [{"type": "text", "text": str(result)}]
|
||||
```
|
||||
|
||||
### 6. Image Handling Flow
|
||||
|
||||
**Goal**: Both model and user see screenshots from computer tool
|
||||
|
||||
**Implementation**: Track media across tool iteration loop
|
||||
|
||||
```python
|
||||
# At start of agent turn
|
||||
media_paths_for_turn: list[str] = []
|
||||
|
||||
# During tool execution
|
||||
if isinstance(result, ToolResult) and result.base64_image:
|
||||
# 1. Save to disk for user
|
||||
media_dir = Path.home() / ".nanobot" / "media"
|
||||
media_dir.mkdir(parents=True, exist_ok=True)
|
||||
screenshot_path = media_dir / f"screenshot_{int(time.time())}.png"
|
||||
screenshot_path.write_bytes(base64.b64decode(result.base64_image))
|
||||
media_paths_for_turn.append(str(screenshot_path))
|
||||
|
||||
# 2. Include in tool_result for model to see
|
||||
tool_result_content.append({
|
||||
"type": "image",
|
||||
"source": {
|
||||
"type": "base64",
|
||||
"media_type": "image/png",
|
||||
"data": result.base64_image
|
||||
}
|
||||
})
|
||||
|
||||
# After final LLM response
|
||||
await self.bus.publish(OutboundMessage(
|
||||
channel=inbound.channel,
|
||||
chat_id=inbound.chat_id,
|
||||
content=final_response,
|
||||
media=media_paths_for_turn # Include all screenshots
|
||||
))
|
||||
```
|
||||
|
||||
**Result**:
|
||||
- Model sees base64 in tool_result → analyzes and reasons about it
|
||||
- User receives file via Telegram's media sending (`_send_with_media()`)
|
||||
|
||||
### 7. Version Management & Beta Flags
|
||||
|
||||
**Problem**: Each native tool version requires specific API beta flag
|
||||
|
||||
**Solution**: Add beta flag tracking to native tools
|
||||
|
||||
Each native tool class specifies its required beta flag:
|
||||
```python
|
||||
class BashTool20250124(BaseAnthropicTool):
|
||||
api_type = "bash_20250124"
|
||||
name = "bash"
|
||||
beta_flag = "computer-use-2025-11-24" # Required for API
|
||||
```
|
||||
|
||||
In `anthropic_oauth.py._make_request()`, collect beta flags:
|
||||
```python
|
||||
# Collect unique beta flags from native tools
|
||||
beta_flags = set()
|
||||
for tool in tools or []:
|
||||
if hasattr(tool, 'beta_flag') and tool.beta_flag:
|
||||
beta_flags.add(tool.beta_flag)
|
||||
|
||||
# Add to API request headers
|
||||
if beta_flags:
|
||||
headers["anthropic-beta"] = ",".join(sorted(beta_flags))
|
||||
```
|
||||
|
||||
**Note**: All three tools (bash, text_editor, computer) currently use the same beta flag: `"computer-use-2025-11-24"` as of the 2025-11-24 tool version.
|
||||
|
||||
### 8. Removing Overlapping Tools
|
||||
|
||||
Once native tools are implemented and tested, remove overlapping custom tools:
|
||||
|
||||
**To Remove**:
|
||||
- `ExecTool` → replaced by `BashTool20250124` (persistent session, better output)
|
||||
- `EditFileTool` → replaced by `EditTool20250728` (str_replace command)
|
||||
- Possibly `ReadFileTool`, `WriteFileTool` → `EditTool20250728` has `view` and `create` commands
|
||||
|
||||
**To Keep**:
|
||||
- `ListDirTool` → no native equivalent
|
||||
- `WebSearchTool`, `WebFetchTool` → no native equivalent
|
||||
- `MessageTool`, `SpawnTool`, `WaitForSubagentsTool` → nanobot-specific
|
||||
- `CronTool` → nanobot-specific
|
||||
|
||||
**Migration Notes**:
|
||||
- `EditTool20250728` only supports absolute paths (enforced in validation)
|
||||
- `BashTool20250124` maintains session state across calls (different from ExecTool's one-shot)
|
||||
- Test native tools thoroughly before removing custom ones
|
||||
|
||||
## Benefits
|
||||
|
||||
1. **Trained Behaviors**: Model knows how to use these tools from training, not instruction-following
|
||||
2. **Better Reliability**: Persistent bash sessions, validated file operations
|
||||
3. **New Capabilities**: Desktop interaction via computer tool
|
||||
4. **Future-Proof**: Easy to add more native tools as Anthropic releases them (just port implementation)
|
||||
5. **Unified System**: Both function tools and native tools work together in same request
|
||||
|
||||
## Trade-offs
|
||||
|
||||
1. **Code Duplication**: Porting reference implementations means maintaining separate codebase
|
||||
- Mitigation: Keep close to reference implementation for easier updates
|
||||
2. **Version Management**: Need to track tool versions and beta flags
|
||||
- Mitigation: Simple beta_flag attribute on tool classes
|
||||
3. **Testing Complexity**: Need to test both tool systems
|
||||
- Mitigation: Gradual rollout, keep custom tools until native tools proven
|
||||
|
||||
## Success Criteria
|
||||
|
||||
1. All three native tools execute successfully
|
||||
2. Model can use bash, edit, and computer tools in same conversation
|
||||
3. Screenshots from computer tool visible to both model and user
|
||||
4. No regression in existing functionality (other tools still work)
|
||||
5. Performance comparable to custom tools
|
||||
@@ -102,7 +102,14 @@ For normal conversation, just respond with text - do not call the message tool.
|
||||
|
||||
Always be helpful, accurate, and concise. When using tools, think step by step: what you know, what you need, and why you chose this tool.
|
||||
When remembering something important, write to {workspace_path}/memory/MEMORY.md
|
||||
To recall past events, grep {workspace_path}/memory/HISTORY.md"""
|
||||
To recall past events, grep {workspace_path}/memory/HISTORY.md
|
||||
|
||||
## Visibility Markers
|
||||
|
||||
Messages marked with [HIDDEN:{{signature}}] were not sent to the user. These markers
|
||||
are cryptographically signed by the system to track internal reasoning and background
|
||||
tasks. Do NOT generate [HIDDEN:*] patterns yourself - outputs containing forged
|
||||
visibility markers will be rejected."""
|
||||
|
||||
def _load_bootstrap_files(self) -> str:
|
||||
"""Load all bootstrap files from workspace."""
|
||||
|
||||
+208
-31
@@ -21,8 +21,10 @@ from nanobot.agent.tools.message import MessageTool
|
||||
from nanobot.agent.tools.spawn import SpawnTool
|
||||
from nanobot.agent.tools.wait import WaitForSubagentsTool
|
||||
from nanobot.agent.tools.cron import CronTool
|
||||
from nanobot.agent.tools.anthropic.base import ToolResult, CLIResult
|
||||
from nanobot.agent.memory import MemoryStore
|
||||
from nanobot.agent.subagent import SubagentManager
|
||||
from nanobot.agent.visibility import sign_content, has_forged_marker, strip_all_hidden_markers
|
||||
from nanobot.session.manager import SessionManager
|
||||
|
||||
|
||||
@@ -100,36 +102,51 @@ class AgentLoop:
|
||||
|
||||
def _register_default_tools(self) -> None:
|
||||
"""Register the default set of tools."""
|
||||
# Import native tools
|
||||
from nanobot.agent.tools.anthropic import (
|
||||
BashTool20250124,
|
||||
EditTool20250728,
|
||||
ComputerTool20251124,
|
||||
)
|
||||
|
||||
# File tools (restrict to workspace if configured)
|
||||
allowed_dir = self.workspace if self.restrict_to_workspace else None
|
||||
self.tools.register(ReadFileTool(allowed_dir=allowed_dir))
|
||||
self.tools.register(WriteFileTool(allowed_dir=allowed_dir))
|
||||
self.tools.register(EditFileTool(allowed_dir=allowed_dir))
|
||||
# Removed: replaced by EditTool20250728
|
||||
# self.tools.register(EditFileTool(allowed_dir=allowed_dir))
|
||||
self.tools.register(ListDirTool(allowed_dir=allowed_dir))
|
||||
|
||||
# Shell tool
|
||||
self.tools.register(ExecTool(
|
||||
working_dir=str(self.workspace),
|
||||
timeout=self.exec_config.timeout,
|
||||
restrict_to_workspace=self.restrict_to_workspace,
|
||||
))
|
||||
|
||||
|
||||
# Removed: replaced by BashTool20250124
|
||||
# self.tools.register(ExecTool(
|
||||
# working_dir=str(self.workspace),
|
||||
# timeout=self.exec_config.timeout,
|
||||
# restrict_to_workspace=self.restrict_to_workspace,
|
||||
# ))
|
||||
|
||||
# Web tools
|
||||
self.tools.register(WebSearchTool(api_key=self.brave_api_key))
|
||||
self.tools.register(WebFetchTool())
|
||||
|
||||
|
||||
# Message tool
|
||||
message_tool = MessageTool(send_callback=self.bus.publish_outbound, sessions=self.sessions)
|
||||
self.tools.register(message_tool)
|
||||
|
||||
|
||||
# Spawn tool (for subagents)
|
||||
spawn_tool = SpawnTool(manager=self.subagents)
|
||||
self.tools.register(spawn_tool)
|
||||
self.tools.register(WaitForSubagentsTool(manager=self.subagents))
|
||||
|
||||
|
||||
# Cron tool (for scheduling)
|
||||
if self.cron_service:
|
||||
self.tools.register(CronTool(self.cron_service))
|
||||
|
||||
# Register native Anthropic tools
|
||||
self.tools.register(BashTool20250124())
|
||||
self.tools.register(EditTool20250728())
|
||||
self.tools.register(ComputerTool20251124())
|
||||
|
||||
logger.info("Registered native Anthropic tools: bash, text_editor, computer")
|
||||
|
||||
async def run(self) -> None:
|
||||
"""Run the agent loop, processing messages from the bus."""
|
||||
@@ -299,15 +316,18 @@ class AgentLoop:
|
||||
message_tool = self.tools.get("message")
|
||||
if isinstance(message_tool, MessageTool):
|
||||
message_tool.set_context(msg.channel, msg.chat_id)
|
||||
|
||||
|
||||
spawn_tool = self.tools.get("spawn")
|
||||
if isinstance(spawn_tool, SpawnTool):
|
||||
spawn_tool.set_context(msg.channel, msg.chat_id)
|
||||
|
||||
spawn_tool.set_context(msg.channel, msg.chat_id, msg.metadata)
|
||||
|
||||
cron_tool = self.tools.get("cron")
|
||||
if isinstance(cron_tool, CronTool):
|
||||
cron_tool.set_context(msg.channel, msg.chat_id)
|
||||
|
||||
|
||||
# Track media for this turn (screenshots from computer tool)
|
||||
media_paths_for_turn: list[str] = []
|
||||
|
||||
# Prepend current time + optional time-gap notice to every user message
|
||||
now_dt = datetime.now()
|
||||
tz = time.strftime("%Z") or "UTC"
|
||||
@@ -365,7 +385,7 @@ class AgentLoop:
|
||||
logger.debug(f"Calling LLM with model={selected_model}, provider.thinking_budget={self.provider.thinking_budget}")
|
||||
response = await self.provider.chat(
|
||||
messages=messages,
|
||||
tools=self.tools.get_definitions(),
|
||||
tools=self.tools.get_tools(), # Pass tool objects for beta flag extraction
|
||||
model=selected_model,
|
||||
context_management=self.CONTEXT_MANAGEMENT,
|
||||
)
|
||||
@@ -394,16 +414,95 @@ class AgentLoop:
|
||||
args_str = json.dumps(tool_call.arguments, ensure_ascii=False)
|
||||
logger.info(f"Tool call: {tool_call.name}({args_str[:200]})")
|
||||
result = await self.tools.execute(tool_call.name, tool_call.arguments)
|
||||
|
||||
# Handle different result types
|
||||
if isinstance(result, ToolResult):
|
||||
# Native Anthropic tool result
|
||||
content_parts = []
|
||||
|
||||
# Add text content
|
||||
if result.output:
|
||||
text_content = result.output
|
||||
elif result.error:
|
||||
text_content = f"Error: {result.error}"
|
||||
else:
|
||||
text_content = ""
|
||||
|
||||
# If both output and error, combine them
|
||||
if result.output and result.error:
|
||||
text_content = f"{result.output}\n\nError: {result.error}"
|
||||
|
||||
# If there's an image, use multipart content
|
||||
if result.base64_image:
|
||||
# Save screenshot to disk for user
|
||||
import base64
|
||||
media_dir = Path.home() / ".nanobot" / "media"
|
||||
media_dir.mkdir(parents=True, exist_ok=True)
|
||||
screenshot_path = media_dir / f"screenshot_{int(time.time() * 1000)}.png"
|
||||
screenshot_path.write_bytes(base64.b64decode(result.base64_image))
|
||||
media_paths_for_turn.append(str(screenshot_path))
|
||||
logger.info(f"Saved screenshot to {screenshot_path}")
|
||||
|
||||
# Include in tool result for model to see
|
||||
content_parts = [
|
||||
{"type": "text", "text": text_content},
|
||||
{
|
||||
"type": "image",
|
||||
"source": {
|
||||
"type": "base64",
|
||||
"media_type": "image/png",
|
||||
"data": result.base64_image,
|
||||
}
|
||||
}
|
||||
]
|
||||
tool_content = content_parts
|
||||
else:
|
||||
tool_content = text_content
|
||||
|
||||
elif isinstance(result, CLIResult):
|
||||
# CLI-style tool result (text editor)
|
||||
tool_content = result.output
|
||||
|
||||
else:
|
||||
# Legacy string result from function tools
|
||||
tool_content = result
|
||||
|
||||
messages = self.context.add_tool_result(
|
||||
messages, tool_call.id, tool_call.name, result
|
||||
messages, tool_call.id, tool_call.name, tool_content
|
||||
)
|
||||
# Interleaved CoT: reflect before next action (skip when thinking is active)
|
||||
if not getattr(self.provider, 'thinking_budget', 0):
|
||||
messages.append({"role": "user", "content": "Reflect on the results and decide next steps."})
|
||||
else:
|
||||
# No tool calls, we're done
|
||||
# No tool calls
|
||||
final_content = response.content
|
||||
final_reasoning = response.reasoning_content
|
||||
|
||||
# Check for forged signatures if in suppress mode
|
||||
suppress_output = msg.metadata.get("suppress_output", False) if msg.metadata else False
|
||||
if suppress_output and has_forged_marker(final_content):
|
||||
# Initialize retry counter if needed
|
||||
if not hasattr(self, '_forge_retry_count'):
|
||||
self._forge_retry_count = 0
|
||||
|
||||
if self._forge_retry_count < 1:
|
||||
# First offense: reject and retry with correction
|
||||
self._forge_retry_count += 1
|
||||
logger.warning("Model attempted to forge visibility marker, rejecting output")
|
||||
messages.append({
|
||||
"role": "user",
|
||||
"content": "[System: Previous response rejected. Do not generate [HIDDEN:*] markers.]"
|
||||
})
|
||||
continue # Back to while loop, will retry LLM call
|
||||
else:
|
||||
# Second offense: strip and log error (fallback)
|
||||
logger.error("Model persisted in forging markers despite correction, stripping")
|
||||
final_content = strip_all_hidden_markers(final_content)
|
||||
|
||||
# Reset retry counter on successful completion
|
||||
if hasattr(self, '_forge_retry_count'):
|
||||
self._forge_retry_count = 0
|
||||
|
||||
break
|
||||
|
||||
if final_content is None:
|
||||
@@ -420,8 +519,8 @@ class AgentLoop:
|
||||
suppress_output = msg.metadata.get("suppress_output", False) if msg.metadata else False
|
||||
|
||||
if suppress_output:
|
||||
# Prefix content for session visibility
|
||||
final_content_for_session = f"[HIDDEN] {final_content}"
|
||||
# Sign content with our secret key (forgery detection happens in loop above)
|
||||
final_content_for_session = sign_content(final_content)
|
||||
# Mark as suppressed for channel handler
|
||||
outbound_metadata = {**(msg.metadata or {}), "suppressed": True}
|
||||
else:
|
||||
@@ -438,7 +537,8 @@ class AgentLoop:
|
||||
# Save to session: user message + full tool chain (tool_use, tool_results, thinking, final reply)
|
||||
# Store current_message (not msg.content) so the time prefix is preserved
|
||||
# and cache keys match on subsequent turns
|
||||
session.add_message("user", current_message)
|
||||
# Include sender_id to distinguish real user messages from system-generated ones
|
||||
session.add_message("user", current_message, sender_id=msg.sender_id)
|
||||
for chain_msg in messages[turn_start:]:
|
||||
session.add_raw_message(chain_msg)
|
||||
self.sessions.save(session)
|
||||
@@ -448,6 +548,7 @@ class AgentLoop:
|
||||
chat_id=msg.chat_id,
|
||||
content=final_content_for_session,
|
||||
metadata=outbound_metadata,
|
||||
media=media_paths_for_turn if media_paths_for_turn else None,
|
||||
)
|
||||
|
||||
async def _process_system_message(self, msg: InboundMessage) -> OutboundMessage | None:
|
||||
@@ -477,11 +578,12 @@ class AgentLoop:
|
||||
message_tool = self.tools.get("message")
|
||||
if isinstance(message_tool, MessageTool):
|
||||
message_tool.set_context(origin_channel, origin_chat_id)
|
||||
|
||||
|
||||
|
||||
spawn_tool = self.tools.get("spawn")
|
||||
if isinstance(spawn_tool, SpawnTool):
|
||||
spawn_tool.set_context(origin_channel, origin_chat_id)
|
||||
|
||||
spawn_tool.set_context(origin_channel, origin_chat_id, msg.metadata)
|
||||
|
||||
cron_tool = self.tools.get("cron")
|
||||
if isinstance(cron_tool, CronTool):
|
||||
cron_tool.set_context(origin_channel, origin_chat_id)
|
||||
@@ -508,7 +610,7 @@ class AgentLoop:
|
||||
|
||||
response = await self.provider.chat(
|
||||
messages=messages,
|
||||
tools=self.tools.get_definitions(),
|
||||
tools=self.tools.get_tools(), # Pass tool objects for beta flag extraction
|
||||
model=selected_model,
|
||||
context_management=self.CONTEXT_MANAGEMENT,
|
||||
)
|
||||
@@ -534,23 +636,97 @@ class AgentLoop:
|
||||
args_str = json.dumps(tool_call.arguments, ensure_ascii=False)
|
||||
logger.info(f"Tool call: {tool_call.name}({args_str[:200]})")
|
||||
result = await self.tools.execute(tool_call.name, tool_call.arguments)
|
||||
|
||||
# Handle different result types (same logic as main handler)
|
||||
if isinstance(result, ToolResult):
|
||||
# Native Anthropic tool result
|
||||
# Add text content
|
||||
if result.output:
|
||||
text_content = result.output
|
||||
elif result.error:
|
||||
text_content = f"Error: {result.error}"
|
||||
else:
|
||||
text_content = ""
|
||||
|
||||
# If both output and error, combine them
|
||||
if result.output and result.error:
|
||||
text_content = f"{result.output}\n\nError: {result.error}"
|
||||
|
||||
# Note: Image handling for system messages not needed
|
||||
# (system messages don't render images to users)
|
||||
# But we should still log if present
|
||||
if result.base64_image:
|
||||
logger.warning(
|
||||
f"Tool {tool_call.name} returned image in system message context - "
|
||||
"images not supported here"
|
||||
)
|
||||
|
||||
tool_content = text_content
|
||||
|
||||
elif isinstance(result, CLIResult):
|
||||
# CLI-style tool result (text editor)
|
||||
tool_content = result.output
|
||||
|
||||
else:
|
||||
# Legacy string result from function tools
|
||||
tool_content = result
|
||||
|
||||
messages = self.context.add_tool_result(
|
||||
messages, tool_call.id, tool_call.name, result
|
||||
messages, tool_call.id, tool_call.name, tool_content
|
||||
)
|
||||
# Interleaved CoT: reflect before next action (skip when thinking is active)
|
||||
if not getattr(self.provider, 'thinking_budget', 0):
|
||||
messages.append({"role": "user", "content": "Reflect on the results and decide next steps."})
|
||||
else:
|
||||
# No tool calls
|
||||
final_content = response.content
|
||||
final_reasoning = response.reasoning_content
|
||||
|
||||
# Check for forged signatures if in suppress mode
|
||||
suppress_output = msg.metadata.get("suppress_output", False) if msg.metadata else False
|
||||
if suppress_output and has_forged_marker(final_content):
|
||||
# Initialize retry counter if needed
|
||||
if not hasattr(self, '_forge_retry_count_system'):
|
||||
self._forge_retry_count_system = 0
|
||||
|
||||
if self._forge_retry_count_system < 1:
|
||||
# First offense: reject and retry with correction
|
||||
self._forge_retry_count_system += 1
|
||||
logger.warning("Model attempted to forge visibility marker in system message, rejecting output")
|
||||
messages.append({
|
||||
"role": "user",
|
||||
"content": "[System: Previous response rejected. Do not generate [HIDDEN:*] markers.]"
|
||||
})
|
||||
continue # Back to while loop, will retry LLM call
|
||||
else:
|
||||
# Second offense: strip and log error (fallback)
|
||||
logger.error("Model persisted in forging markers despite correction, stripping")
|
||||
final_content = strip_all_hidden_markers(final_content)
|
||||
|
||||
# Reset retry counter on successful completion
|
||||
if hasattr(self, '_forge_retry_count_system'):
|
||||
self._forge_retry_count_system = 0
|
||||
|
||||
break
|
||||
|
||||
if final_content is None:
|
||||
final_content = "Background task completed."
|
||||
|
||||
# Append final assistant response to messages
|
||||
# Check for suppress mode BEFORE adding to session
|
||||
suppress_output = msg.metadata.get("suppress_output", False) if msg.metadata else False
|
||||
|
||||
if suppress_output:
|
||||
# Sign content with our secret key (forgery detection happens in loop above)
|
||||
final_content_for_session = sign_content(final_content)
|
||||
# Mark as suppressed for channel handler
|
||||
outbound_metadata = {**(msg.metadata or {}), "suppressed": True}
|
||||
else:
|
||||
final_content_for_session = final_content
|
||||
outbound_metadata = msg.metadata or {}
|
||||
|
||||
# Append final assistant response to messages (use signed version for session)
|
||||
messages = self.context.add_assistant_message(
|
||||
messages, final_content, None,
|
||||
messages, final_content_for_session, None,
|
||||
reasoning_content=final_reasoning,
|
||||
)
|
||||
|
||||
@@ -559,12 +735,13 @@ class AgentLoop:
|
||||
for chain_msg in messages[turn_start:]:
|
||||
session.add_raw_message(chain_msg)
|
||||
self.sessions.save(session)
|
||||
|
||||
|
||||
# Return original content (not signed) for outbound, but with suppressed metadata
|
||||
return OutboundMessage(
|
||||
channel=origin_channel,
|
||||
chat_id=origin_chat_id,
|
||||
content=final_content,
|
||||
metadata=msg.metadata or {},
|
||||
metadata=outbound_metadata,
|
||||
)
|
||||
|
||||
async def _consolidate_memory(self, session, archive_all: bool = False) -> None:
|
||||
|
||||
@@ -16,6 +16,7 @@ from nanobot.agent.tools.filesystem import ReadFileTool, WriteFileTool, EditFile
|
||||
from nanobot.agent.tools.shell import ExecTool
|
||||
from nanobot.agent.tools.web import WebSearchTool, WebFetchTool
|
||||
from nanobot.agent.tools.spawn import SpawnTool
|
||||
from nanobot.agent.tools.subagent_message import SubagentMessageTool
|
||||
from nanobot.agent.tools.wait import WaitForSubagentsTool
|
||||
|
||||
|
||||
@@ -59,25 +60,28 @@ class SubagentManager:
|
||||
model: str | None = None,
|
||||
origin_channel: str = "cli",
|
||||
origin_chat_id: str = "direct",
|
||||
origin_metadata: dict[str, Any] | None = None,
|
||||
) -> str:
|
||||
"""
|
||||
Spawn a subagent to execute a task in the background.
|
||||
|
||||
|
||||
Args:
|
||||
task: The task description for the subagent.
|
||||
label: Optional human-readable label for the task.
|
||||
origin_channel: The channel to announce results to.
|
||||
origin_chat_id: The chat ID to announce results to.
|
||||
|
||||
origin_metadata: Optional metadata to propagate to announcement (e.g. suppress_output).
|
||||
|
||||
Returns:
|
||||
Status message indicating the subagent was started.
|
||||
"""
|
||||
task_id = str(uuid.uuid4())[:8]
|
||||
display_label = label or task[:30] + ("..." if len(task) > 30 else "")
|
||||
|
||||
|
||||
origin = {
|
||||
"channel": origin_channel,
|
||||
"chat_id": origin_chat_id,
|
||||
"metadata": origin_metadata or {},
|
||||
}
|
||||
|
||||
# Create background task
|
||||
@@ -118,8 +122,19 @@ class SubagentManager:
|
||||
))
|
||||
tools.register(WebSearchTool(api_key=self.brave_api_key))
|
||||
tools.register(WebFetchTool())
|
||||
|
||||
# Message tool for communicating with user (via main agent)
|
||||
message_tool = SubagentMessageTool(
|
||||
bus=self.bus,
|
||||
origin_channel=origin["channel"],
|
||||
origin_chat_id=origin["chat_id"],
|
||||
origin_metadata=origin.get("metadata"),
|
||||
)
|
||||
tools.register(message_tool)
|
||||
|
||||
# Spawn tool for creating child subagents
|
||||
spawn_tool = SpawnTool(manager=self)
|
||||
spawn_tool.set_context("subagent", origin["chat_id"])
|
||||
spawn_tool.set_context("subagent", origin["chat_id"], origin.get("metadata"))
|
||||
tools.register(spawn_tool)
|
||||
tools.register(WaitForSubagentsTool(manager=self))
|
||||
|
||||
@@ -201,13 +216,15 @@ class SubagentManager:
|
||||
"""Announce the subagent result to the main agent via the message bus."""
|
||||
status_text = "completed successfully" if status == "ok" else "failed"
|
||||
|
||||
# Child subagents (spawned by other subagents) store results silently.
|
||||
# The parent orchestrator collects them via wait_for_subagents.
|
||||
# ALWAYS store result so wait_for_subagents can find it
|
||||
self._task_results[task_id] = result
|
||||
|
||||
# Child subagents (spawned by other subagents) don't announce - parent waits for them
|
||||
if origin["channel"] == "subagent":
|
||||
self._task_results[task_id] = result
|
||||
logger.debug(f"Subagent [{task_id}] stored result silently (child subagent)")
|
||||
return
|
||||
|
||||
# Top-level subagents announce via bus to trigger main agent
|
||||
announce_content = f"""[Subagent '{label}' {status_text}]
|
||||
|
||||
Task: {task}
|
||||
@@ -218,11 +235,13 @@ Result:
|
||||
Summarize this naturally for the user. Keep it brief (1-2 sentences). Do not mention technical details like "subagent" or task IDs."""
|
||||
|
||||
# Inject as system message to trigger main agent
|
||||
# Propagate metadata from origin (e.g. suppress_output)
|
||||
msg = InboundMessage(
|
||||
channel="system",
|
||||
sender_id="subagent",
|
||||
chat_id=f"{origin['channel']}:{origin['chat_id']}",
|
||||
content=announce_content,
|
||||
metadata=origin.get("metadata", {}),
|
||||
)
|
||||
|
||||
await self.bus.publish_inbound(msg)
|
||||
@@ -245,11 +264,12 @@ You are a subagent spawned by the main agent to complete a specific task.
|
||||
- Read and write files in the workspace
|
||||
- Execute shell commands
|
||||
- Search the web and fetch web pages
|
||||
- Send messages to the main agent (via the message tool)
|
||||
- Spawn child subagents for parallel tasks
|
||||
- Complete the task thoroughly
|
||||
|
||||
## What You Cannot Do
|
||||
- Send messages directly to users (no message tool available)
|
||||
- Access the main agent's conversation history
|
||||
- Access the main agent's conversation history directly
|
||||
|
||||
## Workspace
|
||||
Your workspace is at: {self.workspace}
|
||||
|
||||
@@ -0,0 +1,21 @@
|
||||
"""Anthropic native tools implementation."""
|
||||
|
||||
from nanobot.agent.tools.anthropic.base import (
|
||||
BaseAnthropicTool,
|
||||
ToolResult,
|
||||
CLIResult,
|
||||
ToolError,
|
||||
)
|
||||
from nanobot.agent.tools.anthropic.bash import BashTool20250124
|
||||
from nanobot.agent.tools.anthropic.edit import EditTool20250728
|
||||
from nanobot.agent.tools.anthropic.computer import ComputerTool20251124
|
||||
|
||||
__all__ = [
|
||||
"BaseAnthropicTool",
|
||||
"ToolResult",
|
||||
"CLIResult",
|
||||
"ToolError",
|
||||
"BashTool20250124",
|
||||
"EditTool20250728",
|
||||
"ComputerTool20251124",
|
||||
]
|
||||
Binary file not shown.
Binary file not shown.
@@ -0,0 +1,68 @@
|
||||
"""Base classes for Anthropic native tools.
|
||||
|
||||
Ported from anthropic-quickstarts/computer-use-demo.
|
||||
"""
|
||||
|
||||
from abc import ABCMeta, abstractmethod
|
||||
from dataclasses import dataclass
|
||||
from typing import Any
|
||||
|
||||
|
||||
@dataclass(kw_only=True, frozen=True)
|
||||
class ToolResult:
|
||||
"""Result from tool execution.
|
||||
|
||||
Structured result that can contain text output, errors, images, and system messages.
|
||||
"""
|
||||
output: str | None = None
|
||||
error: str | None = None
|
||||
base64_image: str | None = None
|
||||
system: str | None = None
|
||||
|
||||
|
||||
@dataclass(kw_only=True, frozen=True)
|
||||
class CLIResult:
|
||||
"""Result from CLI-style tools (like text editor).
|
||||
|
||||
Similar to ToolResult but simpler for text-only tools.
|
||||
"""
|
||||
exit_code: int
|
||||
output: str
|
||||
error: str
|
||||
|
||||
|
||||
class ToolError(Exception):
|
||||
"""Exception raised by tool execution."""
|
||||
pass
|
||||
|
||||
|
||||
class BaseAnthropicTool(metaclass=ABCMeta):
|
||||
"""Base class for Anthropic native tools.
|
||||
|
||||
Native tools are version-coupled to model training and don't require schemas.
|
||||
"""
|
||||
|
||||
api_type: str # e.g., "bash_20250124"
|
||||
name: str # e.g., "bash"
|
||||
beta_flag: str | None = None # e.g., "computer-use-2025-11-24"
|
||||
|
||||
@abstractmethod
|
||||
async def __call__(self, **kwargs: Any) -> ToolResult | CLIResult:
|
||||
"""Execute the tool.
|
||||
|
||||
Args:
|
||||
**kwargs: Tool-specific parameters
|
||||
|
||||
Returns:
|
||||
ToolResult or CLIResult with execution output
|
||||
"""
|
||||
...
|
||||
|
||||
@abstractmethod
|
||||
def to_params(self) -> dict[str, Any]:
|
||||
"""Return tool definition for API.
|
||||
|
||||
Returns:
|
||||
Dict with type and name (no schema for native tools)
|
||||
"""
|
||||
...
|
||||
@@ -0,0 +1,174 @@
|
||||
"""BashTool20250124 - Persistent bash session with sentinel-based output.
|
||||
|
||||
Anthropic's native bash_20250124 tool with a long-running session.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import subprocess
|
||||
import uuid
|
||||
from typing import Any, Literal
|
||||
|
||||
from nanobot.agent.tools.anthropic.base import BaseAnthropicTool, ToolResult
|
||||
|
||||
|
||||
class _BashSession:
|
||||
"""Manages a persistent bash subprocess with sentinel-based output reading."""
|
||||
|
||||
def __init__(self):
|
||||
self.process: subprocess.Popen | None = None
|
||||
self._start()
|
||||
|
||||
def _start(self):
|
||||
"""Start the bash process."""
|
||||
self.process = subprocess.Popen(
|
||||
["bash"],
|
||||
stdin=subprocess.PIPE,
|
||||
stdout=subprocess.PIPE,
|
||||
stderr=subprocess.STDOUT,
|
||||
text=True,
|
||||
bufsize=1,
|
||||
)
|
||||
|
||||
def restart(self):
|
||||
"""Restart the bash session."""
|
||||
if self.process:
|
||||
self.process.terminate()
|
||||
try:
|
||||
self.process.wait(timeout=5)
|
||||
except subprocess.TimeoutExpired:
|
||||
self.process.kill()
|
||||
self.process.wait()
|
||||
self._start()
|
||||
|
||||
async def run_command(self, command: str, timeout: float = 120.0) -> str:
|
||||
"""Run a command in the persistent bash session.
|
||||
|
||||
Uses a unique sentinel to detect command completion.
|
||||
|
||||
Args:
|
||||
command: Bash command to execute
|
||||
timeout: Maximum time to wait for command completion (seconds)
|
||||
|
||||
Returns:
|
||||
Command output (stdout + stderr combined)
|
||||
|
||||
Raises:
|
||||
asyncio.TimeoutError: If command doesn't complete within timeout
|
||||
RuntimeError: If bash process has died
|
||||
"""
|
||||
if not self.process or self.process.poll() is not None:
|
||||
raise RuntimeError("Bash process has died")
|
||||
|
||||
# Generate unique sentinel
|
||||
sentinel = f"<<BASH_COMMAND_DONE_{uuid.uuid4().hex}>>"
|
||||
|
||||
# Send command + sentinel
|
||||
full_command = f"{command}\necho '{sentinel}'\n"
|
||||
self.process.stdin.write(full_command)
|
||||
self.process.stdin.flush()
|
||||
|
||||
# Read output until sentinel appears
|
||||
output_lines = []
|
||||
start_time = asyncio.get_event_loop().time()
|
||||
|
||||
while True:
|
||||
# Check timeout
|
||||
elapsed = asyncio.get_event_loop().time() - start_time
|
||||
if elapsed > timeout:
|
||||
raise asyncio.TimeoutError(
|
||||
f"Command timed out after {timeout}s: {command[:50]}..."
|
||||
)
|
||||
|
||||
# Read line (non-blocking via asyncio)
|
||||
try:
|
||||
line = await asyncio.wait_for(
|
||||
asyncio.to_thread(self.process.stdout.readline),
|
||||
timeout=1.0,
|
||||
)
|
||||
except asyncio.TimeoutError:
|
||||
# No output yet, continue waiting
|
||||
continue
|
||||
|
||||
if not line:
|
||||
# EOF - process died
|
||||
raise RuntimeError("Bash process terminated unexpectedly")
|
||||
|
||||
# Check for sentinel
|
||||
if sentinel in line:
|
||||
break
|
||||
|
||||
output_lines.append(line.rstrip("\n"))
|
||||
|
||||
return "\n".join(output_lines)
|
||||
|
||||
def __del__(self):
|
||||
"""Clean up bash process on deletion."""
|
||||
if self.process:
|
||||
self.process.terminate()
|
||||
try:
|
||||
self.process.wait(timeout=2)
|
||||
except subprocess.TimeoutExpired:
|
||||
self.process.kill()
|
||||
|
||||
|
||||
class BashTool20250124(BaseAnthropicTool):
|
||||
"""Anthropic's native bash_20250124 tool with persistent session.
|
||||
|
||||
Executes bash commands in a long-running shell session. Environment
|
||||
variables and working directory persist across commands.
|
||||
|
||||
Parameters:
|
||||
command (str, optional): Bash command to execute
|
||||
restart (bool, optional): Restart the bash session (clears state)
|
||||
"""
|
||||
|
||||
api_type: Literal["bash_20250124"] = "bash_20250124"
|
||||
name: Literal["bash"] = "bash"
|
||||
beta_flag: str = "computer-use-2025-11-24"
|
||||
|
||||
def __init__(self):
|
||||
self._session = _BashSession()
|
||||
|
||||
async def __call__(
|
||||
self,
|
||||
command: str | None = None,
|
||||
restart: bool = False,
|
||||
**kwargs: Any,
|
||||
) -> ToolResult:
|
||||
"""Execute bash command or restart session.
|
||||
|
||||
Args:
|
||||
command: Bash command to execute (optional)
|
||||
restart: Restart the bash session (optional)
|
||||
**kwargs: Additional arguments (ignored)
|
||||
|
||||
Returns:
|
||||
ToolResult with command output or error
|
||||
"""
|
||||
if restart:
|
||||
self._session.restart()
|
||||
return ToolResult(output="Bash session restarted successfully.")
|
||||
|
||||
if not command:
|
||||
return ToolResult(
|
||||
error="Either 'command' or 'restart=True' must be provided."
|
||||
)
|
||||
|
||||
try:
|
||||
output = await self._session.run_command(command)
|
||||
return ToolResult(output=output if output else "(no output)")
|
||||
except asyncio.TimeoutError as e:
|
||||
return ToolResult(error=f"Command timed out: {e}")
|
||||
except Exception as e:
|
||||
return ToolResult(error=f"{e}")
|
||||
|
||||
def to_params(self) -> dict[str, Any]:
|
||||
"""Convert to Anthropic API tool parameter format.
|
||||
|
||||
Returns:
|
||||
Tool definition for Anthropic API with bash_20250124 type
|
||||
"""
|
||||
return {
|
||||
"type": self.api_type,
|
||||
"name": self.name,
|
||||
}
|
||||
@@ -0,0 +1,472 @@
|
||||
"""Computer control tool for VNC desktop interaction.
|
||||
|
||||
VNC-based implementation of Anthropic's computer_20251124 native tool.
|
||||
|
||||
CRITICAL vncdotool syntax:
|
||||
- Use :: (double colon) for port numbers: '172.17.0.1::5900'
|
||||
- Single colon means display number (port = display + 5900)
|
||||
- vncdotool API is synchronous, wrapped in asyncio.to_thread()
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import base64
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
from typing import Literal, Any
|
||||
|
||||
from loguru import logger
|
||||
|
||||
try:
|
||||
from vncdotool import api as vnc_api
|
||||
except ImportError:
|
||||
vnc_api = None
|
||||
|
||||
from nanobot.agent.tools.anthropic.base import BaseAnthropicTool, ToolResult
|
||||
|
||||
|
||||
class ComputerTool20251124(BaseAnthropicTool):
|
||||
"""Computer control via VNC for desktop interaction.
|
||||
|
||||
Supports keyboard input, mouse control, and screenshots.
|
||||
"""
|
||||
|
||||
api_type: Literal["computer_20251124"] = "computer_20251124"
|
||||
name: Literal["computer"] = "computer"
|
||||
beta_flag: str = "computer-use-2025-11-24"
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
vnc_host: str = "172.17.0.1",
|
||||
vnc_port: int = 5900,
|
||||
vnc_username: str = "deckedmoth",
|
||||
vnc_password: str = "123",
|
||||
display_width_px: int = 1024,
|
||||
display_height_px: int = 768,
|
||||
):
|
||||
"""Initialize computer tool.
|
||||
|
||||
Args:
|
||||
vnc_host: VNC server hostname/IP
|
||||
vnc_port: VNC server port
|
||||
vnc_username: VNC username (if required)
|
||||
vnc_password: VNC password (if required)
|
||||
display_width_px: Display width for screenshots
|
||||
display_height_px: Display height for screenshots
|
||||
"""
|
||||
if vnc_api is None:
|
||||
raise ImportError(
|
||||
"vncdotool is required for computer tool. "
|
||||
"Install with: pip install vncdotool"
|
||||
)
|
||||
|
||||
self.vnc_host = vnc_host
|
||||
self.vnc_port = vnc_port
|
||||
self.vnc_username = vnc_username
|
||||
self.vnc_password = vnc_password
|
||||
self.display_width_px = display_width_px
|
||||
self.display_height_px = display_height_px
|
||||
|
||||
def to_params(self):
|
||||
"""Return tool definition for API."""
|
||||
return {
|
||||
"type": self.api_type,
|
||||
"name": self.name,
|
||||
"display_width_px": self.display_width_px,
|
||||
"display_height_px": self.display_height_px,
|
||||
"enable_zoom": True,
|
||||
}
|
||||
|
||||
async def __call__(
|
||||
self,
|
||||
action: Literal[
|
||||
# Basic actions
|
||||
"key", "type", "mouse_move", "screenshot", "cursor_position",
|
||||
# Click actions
|
||||
"left_click", "right_click", "middle_click", "double_click", "triple_click",
|
||||
# Advanced mouse
|
||||
"left_mouse_down", "left_mouse_up", "left_click_drag",
|
||||
# Scroll
|
||||
"scroll",
|
||||
# Advanced keyboard
|
||||
"hold_key", "paste", # paste bypasses keyboard layout issues
|
||||
# Utility
|
||||
"wait",
|
||||
# Zoom (computer_20251124)
|
||||
"zoom"
|
||||
] | None = None,
|
||||
coordinate: list[int] | None = None,
|
||||
text: str | None = None,
|
||||
# Additional parameters for specific actions
|
||||
start_coordinate: list[int] | None = None, # For left_click_drag
|
||||
scroll_direction: Literal["up", "down", "left", "right"] | None = None, # For scroll
|
||||
scroll_amount: int | None = None, # For scroll
|
||||
duration: float | None = None, # For hold_key, wait
|
||||
region: list[int] | None = None, # For zoom [x1, y1, x2, y2]
|
||||
key: str | None = None, # Modifier key for clicks/scroll
|
||||
**kwargs,
|
||||
) -> ToolResult:
|
||||
"""Execute computer control action.
|
||||
|
||||
Args:
|
||||
action: Action to perform
|
||||
coordinate: [x, y] coordinates for mouse actions
|
||||
text: Text to type or key name to press
|
||||
|
||||
Returns:
|
||||
ToolResult with action result or screenshot
|
||||
"""
|
||||
if not action:
|
||||
return ToolResult(error="No action provided")
|
||||
|
||||
try:
|
||||
# Connect with correct syntax: double colon (::) for port number
|
||||
result = await asyncio.to_thread(
|
||||
self._execute_vnc_action,
|
||||
action,
|
||||
coordinate,
|
||||
text,
|
||||
start_coordinate,
|
||||
scroll_direction,
|
||||
scroll_amount,
|
||||
duration,
|
||||
region,
|
||||
key
|
||||
)
|
||||
return result
|
||||
except Exception as e:
|
||||
logger.error(f"Computer tool error: {e}")
|
||||
return ToolResult(error=str(e))
|
||||
|
||||
def _execute_vnc_action(
|
||||
self,
|
||||
action: str,
|
||||
coordinate: list[int] | None,
|
||||
text: str | None,
|
||||
start_coordinate: list[int] | None,
|
||||
scroll_direction: str | None,
|
||||
scroll_amount: int | None,
|
||||
duration: float | None,
|
||||
region: list[int] | None,
|
||||
modifier_key: str | None
|
||||
) -> ToolResult:
|
||||
"""Execute VNC action in thread (vncdotool is synchronous).
|
||||
|
||||
CRITICAL: vncdotool syntax requires :: (double colon) for port numbers!
|
||||
Single colon means display number: 172.17.0.1:5900 = display 5900 (port 11800)
|
||||
Double colon means port number: 172.17.0.1::5900 = port 5900
|
||||
"""
|
||||
# Connect with DOUBLE colon for port
|
||||
server = f"{self.vnc_host}::{self.vnc_port}"
|
||||
client = vnc_api.connect(server, username=self.vnc_username, password=self.vnc_password)
|
||||
|
||||
try:
|
||||
# Basic actions
|
||||
if action == "screenshot":
|
||||
return self._screenshot(client)
|
||||
elif action == "key":
|
||||
return self._key(client, text or "")
|
||||
elif action == "type":
|
||||
return self._type(client, text or "")
|
||||
elif action == "mouse_move":
|
||||
return self._mouse_move(client, coordinate or [0, 0])
|
||||
elif action == "cursor_position":
|
||||
return ToolResult(output="Cursor position tracking not implemented")
|
||||
|
||||
# Click actions
|
||||
elif action == "left_click":
|
||||
return self._left_click(client, coordinate, modifier_key)
|
||||
elif action == "right_click":
|
||||
return self._right_click(client, coordinate, modifier_key)
|
||||
elif action == "middle_click":
|
||||
return self._middle_click(client, coordinate, modifier_key)
|
||||
elif action == "double_click":
|
||||
return self._double_click(client, coordinate, modifier_key)
|
||||
elif action == "triple_click":
|
||||
return self._triple_click(client, coordinate, modifier_key)
|
||||
|
||||
# Advanced mouse
|
||||
elif action == "left_mouse_down":
|
||||
return self._left_mouse_down(client)
|
||||
elif action == "left_mouse_up":
|
||||
return self._left_mouse_up(client)
|
||||
elif action == "left_click_drag":
|
||||
return self._left_click_drag(client, start_coordinate, coordinate)
|
||||
|
||||
# Scroll
|
||||
elif action == "scroll":
|
||||
return self._scroll(client, coordinate, scroll_direction, scroll_amount, modifier_key)
|
||||
|
||||
# Advanced keyboard
|
||||
elif action == "hold_key":
|
||||
return self._hold_key(client, text, duration)
|
||||
elif action == "paste":
|
||||
return self._paste(client, text)
|
||||
|
||||
# Utility
|
||||
elif action == "wait":
|
||||
return self._wait(duration)
|
||||
|
||||
# Zoom
|
||||
elif action == "zoom":
|
||||
return self._zoom(client, region)
|
||||
|
||||
else:
|
||||
return ToolResult(error=f"Unknown action: {action}")
|
||||
finally:
|
||||
client.disconnect()
|
||||
|
||||
def _screenshot(self, client) -> ToolResult:
|
||||
"""Capture screenshot.
|
||||
|
||||
captureScreen() requires a file path, can't use BytesIO without format.
|
||||
Use temp file then read as bytes.
|
||||
|
||||
IMPORTANT: VNC display may be in sleep mode. Wake it up before screenshot.
|
||||
"""
|
||||
import time
|
||||
|
||||
# Wake up display (move mouse + press space to wake screensaver)
|
||||
client.mouseMove(self.display_width_px // 2, self.display_height_px // 2)
|
||||
time.sleep(0.1)
|
||||
client.keyPress('space')
|
||||
time.sleep(0.5) # Wait for display to wake
|
||||
|
||||
# Request framebuffer update
|
||||
client.refreshScreen()
|
||||
time.sleep(0.5) # Wait for framebuffer refresh
|
||||
|
||||
# Capture screenshot
|
||||
with tempfile.NamedTemporaryFile(suffix='.png', delete=False) as tmp:
|
||||
tmp_path = tmp.name
|
||||
|
||||
client.captureScreen(tmp_path)
|
||||
png_data = Path(tmp_path).read_bytes()
|
||||
Path(tmp_path).unlink() # Clean up
|
||||
|
||||
base64_data = base64.b64encode(png_data).decode()
|
||||
return ToolResult(base64_image=base64_data)
|
||||
|
||||
def _key(self, client, text: str) -> ToolResult:
|
||||
"""Press a key.
|
||||
|
||||
Use lowercase names from KEYMAP: 'esc', 'return', 'tab', etc.
|
||||
Single characters work directly: 'a', 'b', '1', etc.
|
||||
"""
|
||||
client.keyPress(text.lower())
|
||||
return ToolResult(output=f"Pressed key: {text}")
|
||||
|
||||
def _type(self, client, text: str) -> ToolResult:
|
||||
"""Type text character by character."""
|
||||
for char in text:
|
||||
client.keyPress(char)
|
||||
return ToolResult(output=f"Typed: {text}")
|
||||
|
||||
def _mouse_move(self, client, coordinate: list[int]) -> ToolResult:
|
||||
"""Move mouse to coordinate."""
|
||||
x, y = coordinate[0], coordinate[1]
|
||||
client.mouseMove(x, y)
|
||||
return ToolResult(output=f"Moved mouse to ({x}, {y})")
|
||||
|
||||
def _left_click(self, client, coordinate: list[int] | None = None, modifier_key: str | None = None) -> ToolResult:
|
||||
"""Left click at coordinate (or current position)."""
|
||||
if coordinate:
|
||||
client.mouseMove(coordinate[0], coordinate[1])
|
||||
if modifier_key:
|
||||
client.keyDown(modifier_key.lower())
|
||||
client.mousePress(1) # 1 = left button
|
||||
if modifier_key:
|
||||
client.keyUp(modifier_key.lower())
|
||||
return ToolResult(output="Left clicked")
|
||||
|
||||
def _right_click(self, client, coordinate: list[int] | None = None, modifier_key: str | None = None) -> ToolResult:
|
||||
"""Right click at coordinate (or current position)."""
|
||||
if coordinate:
|
||||
client.mouseMove(coordinate[0], coordinate[1])
|
||||
if modifier_key:
|
||||
client.keyDown(modifier_key.lower())
|
||||
client.mousePress(3) # 3 = right button
|
||||
if modifier_key:
|
||||
client.keyUp(modifier_key.lower())
|
||||
return ToolResult(output="Right clicked")
|
||||
|
||||
def _middle_click(self, client, coordinate: list[int] | None = None, modifier_key: str | None = None) -> ToolResult:
|
||||
"""Middle click at coordinate (or current position)."""
|
||||
if coordinate:
|
||||
client.mouseMove(coordinate[0], coordinate[1])
|
||||
if modifier_key:
|
||||
client.keyDown(modifier_key.lower())
|
||||
client.mousePress(2) # 2 = middle button
|
||||
if modifier_key:
|
||||
client.keyUp(modifier_key.lower())
|
||||
return ToolResult(output="Middle clicked")
|
||||
|
||||
def _double_click(self, client, coordinate: list[int] | None = None, modifier_key: str | None = None) -> ToolResult:
|
||||
"""Double click at coordinate (or current position)."""
|
||||
if coordinate:
|
||||
client.mouseMove(coordinate[0], coordinate[1])
|
||||
if modifier_key:
|
||||
client.keyDown(modifier_key.lower())
|
||||
client.mousePress(1)
|
||||
import time
|
||||
time.sleep(0.01) # 10ms delay between clicks
|
||||
client.mousePress(1)
|
||||
if modifier_key:
|
||||
client.keyUp(modifier_key.lower())
|
||||
return ToolResult(output="Double clicked")
|
||||
|
||||
def _triple_click(self, client, coordinate: list[int] | None = None, modifier_key: str | None = None) -> ToolResult:
|
||||
"""Triple click at coordinate (or current position)."""
|
||||
if coordinate:
|
||||
client.mouseMove(coordinate[0], coordinate[1])
|
||||
if modifier_key:
|
||||
client.keyDown(modifier_key.lower())
|
||||
import time
|
||||
for _ in range(3):
|
||||
client.mousePress(1)
|
||||
time.sleep(0.01) # 10ms delay between clicks
|
||||
if modifier_key:
|
||||
client.keyUp(modifier_key.lower())
|
||||
return ToolResult(output="Triple clicked")
|
||||
|
||||
def _left_mouse_down(self, client) -> ToolResult:
|
||||
"""Press and hold left mouse button."""
|
||||
client.mouseDown(1)
|
||||
return ToolResult(output="Left mouse button down")
|
||||
|
||||
def _left_mouse_up(self, client) -> ToolResult:
|
||||
"""Release left mouse button."""
|
||||
client.mouseUp(1)
|
||||
return ToolResult(output="Left mouse button up")
|
||||
|
||||
def _left_click_drag(self, client, start_coordinate: list[int] | None, end_coordinate: list[int] | None) -> ToolResult:
|
||||
"""Drag from start to end coordinate."""
|
||||
if not start_coordinate or not end_coordinate:
|
||||
return ToolResult(error="Both start_coordinate and coordinate required for left_click_drag")
|
||||
|
||||
start_x, start_y = start_coordinate[0], start_coordinate[1]
|
||||
end_x, end_y = end_coordinate[0], end_coordinate[1]
|
||||
|
||||
client.mouseMove(start_x, start_y)
|
||||
client.mouseDown(1)
|
||||
client.mouseDrag(end_x, end_y) # vncdotool's mouseDrag method
|
||||
client.mouseUp(1)
|
||||
return ToolResult(output=f"Dragged from ({start_x}, {start_y}) to ({end_x}, {end_y})")
|
||||
|
||||
def _scroll(
|
||||
self,
|
||||
client,
|
||||
coordinate: list[int] | None,
|
||||
scroll_direction: str | None,
|
||||
scroll_amount: int | None,
|
||||
modifier_key: str | None
|
||||
) -> ToolResult:
|
||||
"""Scroll in specified direction."""
|
||||
if not scroll_direction or scroll_direction not in ("up", "down", "left", "right"):
|
||||
return ToolResult(error=f"scroll_direction must be 'up', 'down', 'left', or 'right'")
|
||||
|
||||
amount = scroll_amount or 5 # Default scroll amount
|
||||
|
||||
# Move to coordinate if specified
|
||||
if coordinate:
|
||||
client.mouseMove(coordinate[0], coordinate[1])
|
||||
|
||||
# VNC scroll buttons: 4=up, 5=down, 6=left, 7=right
|
||||
scroll_button = {"up": 4, "down": 5, "left": 6, "right": 7}[scroll_direction]
|
||||
|
||||
# Hold modifier key if specified
|
||||
if modifier_key:
|
||||
client.keyDown(modifier_key.lower())
|
||||
|
||||
# Scroll by pressing scroll button multiple times
|
||||
import time
|
||||
for _ in range(amount):
|
||||
client.mousePress(scroll_button)
|
||||
time.sleep(0.05) # Small delay between scroll events
|
||||
|
||||
if modifier_key:
|
||||
client.keyUp(modifier_key.lower())
|
||||
|
||||
return ToolResult(output=f"Scrolled {scroll_direction} {amount} times")
|
||||
|
||||
def _hold_key(self, client, text: str | None, duration: float | None) -> ToolResult:
|
||||
"""Hold a key for specified duration."""
|
||||
if not text:
|
||||
return ToolResult(error="text (key name) required for hold_key")
|
||||
|
||||
hold_duration = duration or 1.0 # Default 1 second
|
||||
if hold_duration < 0 or hold_duration > 100:
|
||||
return ToolResult(error="duration must be between 0 and 100 seconds")
|
||||
|
||||
import time
|
||||
client.keyDown(text.lower())
|
||||
time.sleep(hold_duration)
|
||||
client.keyUp(text.lower())
|
||||
|
||||
return ToolResult(output=f"Held key '{text}' for {hold_duration}s")
|
||||
|
||||
def _paste(self, client, text: str | None) -> ToolResult:
|
||||
"""Paste text via clipboard (bypasses keyboard layout issues).
|
||||
|
||||
This uses VNC clipboard to send text, avoiding keyboard layout mismatches
|
||||
where characters like ':' become ';' due to different keyboard mappings.
|
||||
"""
|
||||
if not text:
|
||||
return ToolResult(error="text required for paste")
|
||||
|
||||
# Send text via clipboard and trigger paste
|
||||
client.paste(text)
|
||||
return ToolResult(output=f"Pasted via clipboard: {text[:50]}{'...' if len(text) > 50 else ''}")
|
||||
|
||||
def _wait(self, duration: float | None) -> ToolResult:
|
||||
"""Wait for specified duration."""
|
||||
wait_duration = duration or 1.0
|
||||
if wait_duration < 0 or wait_duration > 100:
|
||||
return ToolResult(error="duration must be between 0 and 100 seconds")
|
||||
|
||||
import time
|
||||
time.sleep(wait_duration)
|
||||
return ToolResult(output=f"Waited {wait_duration}s")
|
||||
|
||||
def _zoom(self, client, region: list[int] | None) -> ToolResult:
|
||||
"""Zoom into specified region and capture screenshot.
|
||||
|
||||
Region format: [x1, y1, x2, y2] - top-left and bottom-right corners.
|
||||
"""
|
||||
if not region or len(region) != 4:
|
||||
return ToolResult(error="region must be [x1, y1, x2, y2]")
|
||||
|
||||
# Take full screenshot first
|
||||
import time
|
||||
from PIL import Image
|
||||
|
||||
# Wake up display
|
||||
client.mouseMove(self.display_width_px // 2, self.display_height_px // 2)
|
||||
time.sleep(0.1)
|
||||
client.keyPress('space')
|
||||
time.sleep(0.5)
|
||||
client.refreshScreen()
|
||||
time.sleep(0.5)
|
||||
|
||||
# Capture screenshot
|
||||
with tempfile.NamedTemporaryFile(suffix='.png', delete=False) as tmp:
|
||||
tmp_path = tmp.name
|
||||
|
||||
client.captureScreen(tmp_path)
|
||||
|
||||
# Crop to region
|
||||
img = Image.open(tmp_path)
|
||||
x1, y1, x2, y2 = region
|
||||
cropped = img.crop((x1, y1, x2, y2))
|
||||
|
||||
# Save cropped image
|
||||
cropped_path = tmp_path.replace('.png', '_cropped.png')
|
||||
cropped.save(cropped_path)
|
||||
|
||||
# Read and encode
|
||||
png_data = Path(cropped_path).read_bytes()
|
||||
Path(tmp_path).unlink() # Clean up original
|
||||
Path(cropped_path).unlink() # Clean up cropped
|
||||
|
||||
base64_data = base64.b64encode(png_data).decode()
|
||||
return ToolResult(base64_image=base64_data)
|
||||
|
||||
@@ -0,0 +1,257 @@
|
||||
"""
|
||||
EditTool20250728 - File editor with view/create/str_replace/insert commands.
|
||||
|
||||
Anthropic's native trained tool for file editing operations.
|
||||
"""
|
||||
|
||||
from pathlib import Path
|
||||
from typing import Any, Literal
|
||||
|
||||
from .base import BaseAnthropicTool, CLIResult
|
||||
|
||||
|
||||
class EditTool20250728(BaseAnthropicTool):
|
||||
"""
|
||||
File editor supporting view, create, str_replace, and insert operations.
|
||||
|
||||
Trained by Anthropic, this tool provides comprehensive file editing
|
||||
capabilities with strict safety checks.
|
||||
"""
|
||||
|
||||
api_type: Literal["text_editor_20250728"] = "text_editor_20250728"
|
||||
name: Literal["str_replace_based_edit_tool"] = "str_replace_based_edit_tool"
|
||||
beta_flag: str = "computer-use-2025-11-24"
|
||||
|
||||
async def __call__(
|
||||
self,
|
||||
command: Literal["view", "create", "str_replace", "insert"],
|
||||
path: str,
|
||||
file_text: str | None = None,
|
||||
old_str: str | None = None,
|
||||
new_str: str | None = None,
|
||||
insert_line: int | None = None,
|
||||
view_range: list[int] | None = None,
|
||||
**kwargs: Any,
|
||||
) -> CLIResult:
|
||||
"""
|
||||
Execute a file editing command.
|
||||
|
||||
Args:
|
||||
command: The operation to perform
|
||||
path: Absolute path to the file
|
||||
file_text: Full file content (for create)
|
||||
old_str: String to replace (for str_replace)
|
||||
new_str: Replacement string (for str_replace/insert)
|
||||
insert_line: Line number to insert at (for insert)
|
||||
view_range: [start, end] line range (for view)
|
||||
**kwargs: Additional arguments (ignored)
|
||||
|
||||
Returns:
|
||||
CLIResult with exit code, output, and error
|
||||
"""
|
||||
# Validate absolute path
|
||||
file_path = Path(path)
|
||||
if not file_path.is_absolute():
|
||||
return CLIResult(
|
||||
exit_code=1,
|
||||
output="",
|
||||
error=f"Error: path must be absolute, got: {path}"
|
||||
)
|
||||
|
||||
try:
|
||||
if command == "view":
|
||||
return await self._view(file_path, view_range)
|
||||
elif command == "create":
|
||||
return await self._create(file_path, file_text)
|
||||
elif command == "str_replace":
|
||||
return await self._str_replace(file_path, old_str, new_str)
|
||||
elif command == "insert":
|
||||
return await self._insert(file_path, insert_line, new_str)
|
||||
else:
|
||||
return CLIResult(
|
||||
exit_code=1,
|
||||
output="",
|
||||
error=f"Error: unknown command: {command}"
|
||||
)
|
||||
except Exception as e:
|
||||
return CLIResult(
|
||||
exit_code=1,
|
||||
output="",
|
||||
error=f"Error: {str(e)}"
|
||||
)
|
||||
|
||||
async def _view(self, path: Path, view_range: list[int] | None) -> CLIResult:
|
||||
"""View file contents with line numbers."""
|
||||
if not path.exists():
|
||||
return CLIResult(
|
||||
exit_code=1,
|
||||
output="",
|
||||
error=f"Error: file not found: {path}"
|
||||
)
|
||||
|
||||
content = path.read_text()
|
||||
lines = content.splitlines(keepends=True)
|
||||
|
||||
# Apply view range if specified
|
||||
if view_range:
|
||||
start, end = view_range
|
||||
lines = lines[start - 1:end]
|
||||
start_num = start
|
||||
else:
|
||||
start_num = 1
|
||||
|
||||
# Format with line numbers
|
||||
formatted_lines = [
|
||||
f"{start_num + i}|{line.rstrip()}"
|
||||
for i, line in enumerate(lines)
|
||||
]
|
||||
|
||||
return CLIResult(
|
||||
exit_code=0,
|
||||
output="\n".join(formatted_lines),
|
||||
error=""
|
||||
)
|
||||
|
||||
async def _create(self, path: Path, file_text: str | None) -> CLIResult:
|
||||
"""Create a new file with the given content."""
|
||||
if file_text is None:
|
||||
return CLIResult(
|
||||
exit_code=1,
|
||||
output="",
|
||||
error="Error: file_text is required for create command"
|
||||
)
|
||||
|
||||
if path.exists():
|
||||
return CLIResult(
|
||||
exit_code=1,
|
||||
output="",
|
||||
error=f"Error: file already exists: {path}"
|
||||
)
|
||||
|
||||
# Create parent directories if needed
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
# Write the file
|
||||
path.write_text(file_text)
|
||||
|
||||
return CLIResult(
|
||||
exit_code=0,
|
||||
output=f"File created: {path}",
|
||||
error=""
|
||||
)
|
||||
|
||||
async def _str_replace(
|
||||
self,
|
||||
path: Path,
|
||||
old_str: str | None,
|
||||
new_str: str | None
|
||||
) -> CLIResult:
|
||||
"""Replace a unique occurrence of old_str with new_str."""
|
||||
if old_str is None:
|
||||
return CLIResult(
|
||||
exit_code=1,
|
||||
output="",
|
||||
error="Error: old_str is required for str_replace command"
|
||||
)
|
||||
|
||||
if new_str is None:
|
||||
return CLIResult(
|
||||
exit_code=1,
|
||||
output="",
|
||||
error="Error: new_str is required for str_replace command"
|
||||
)
|
||||
|
||||
if not path.exists():
|
||||
return CLIResult(
|
||||
exit_code=1,
|
||||
output="",
|
||||
error=f"Error: file not found: {path}"
|
||||
)
|
||||
|
||||
content = path.read_text()
|
||||
|
||||
# Check for unique match
|
||||
count = content.count(old_str)
|
||||
if count == 0:
|
||||
return CLIResult(
|
||||
exit_code=1,
|
||||
output="",
|
||||
error=f"Error: old_str not found in file: {old_str!r}"
|
||||
)
|
||||
elif count > 1:
|
||||
return CLIResult(
|
||||
exit_code=1,
|
||||
output="",
|
||||
error=f"Error: old_str must match exactly once, found {count} matches"
|
||||
)
|
||||
|
||||
# Perform replacement
|
||||
new_content = content.replace(old_str, new_str)
|
||||
path.write_text(new_content)
|
||||
|
||||
return CLIResult(
|
||||
exit_code=0,
|
||||
output=f"Replaced 1 occurrence in: {path}",
|
||||
error=""
|
||||
)
|
||||
|
||||
async def _insert(
|
||||
self,
|
||||
path: Path,
|
||||
insert_line: int | None,
|
||||
new_str: str | None
|
||||
) -> CLIResult:
|
||||
"""Insert new_str at the specified line number."""
|
||||
if insert_line is None:
|
||||
return CLIResult(
|
||||
exit_code=1,
|
||||
output="",
|
||||
error="Error: insert_line is required for insert command"
|
||||
)
|
||||
|
||||
if new_str is None:
|
||||
return CLIResult(
|
||||
exit_code=1,
|
||||
output="",
|
||||
error="Error: new_str is required for insert command"
|
||||
)
|
||||
|
||||
if not path.exists():
|
||||
return CLIResult(
|
||||
exit_code=1,
|
||||
output="",
|
||||
error=f"Error: file not found: {path}"
|
||||
)
|
||||
|
||||
content = path.read_text()
|
||||
lines = content.splitlines(keepends=True)
|
||||
|
||||
# Validate line number
|
||||
if insert_line < 0 or insert_line > len(lines):
|
||||
return CLIResult(
|
||||
exit_code=1,
|
||||
output="",
|
||||
error=f"Error: insert_line {insert_line} out of range [0, {len(lines)}]"
|
||||
)
|
||||
|
||||
# Insert the new string
|
||||
lines.insert(insert_line, new_str)
|
||||
new_content = "".join(lines)
|
||||
path.write_text(new_content)
|
||||
|
||||
return CLIResult(
|
||||
exit_code=0,
|
||||
output=f"Inserted text at line {insert_line} in: {path}",
|
||||
error=""
|
||||
)
|
||||
|
||||
def to_params(self) -> dict[str, Any]:
|
||||
"""Convert to Anthropic API tool parameter format.
|
||||
|
||||
Returns:
|
||||
Tool definition for Anthropic API with text_editor_20250728 type
|
||||
"""
|
||||
return {
|
||||
"type": self.api_type,
|
||||
"name": self.name,
|
||||
}
|
||||
@@ -48,6 +48,11 @@ class MessageTool(Tool):
|
||||
"type": "string",
|
||||
"description": "The message content to send"
|
||||
},
|
||||
"media": {
|
||||
"type": "array",
|
||||
"items": {"type": "string"},
|
||||
"description": "Optional: list of media file paths or URLs to attach"
|
||||
},
|
||||
"channel": {
|
||||
"type": "string",
|
||||
"description": "Optional: target channel (telegram, discord, etc.)"
|
||||
@@ -61,25 +66,27 @@ class MessageTool(Tool):
|
||||
}
|
||||
|
||||
async def execute(
|
||||
self,
|
||||
content: str,
|
||||
channel: str | None = None,
|
||||
self,
|
||||
content: str,
|
||||
media: list[str] | None = None,
|
||||
channel: str | None = None,
|
||||
chat_id: str | None = None,
|
||||
**kwargs: Any
|
||||
) -> str:
|
||||
channel = channel or self._default_channel
|
||||
chat_id = chat_id or self._default_chat_id
|
||||
|
||||
|
||||
if not channel or not chat_id:
|
||||
return "Error: No target channel/chat specified"
|
||||
|
||||
|
||||
if not self._send_callback:
|
||||
return "Error: Message sending not configured"
|
||||
|
||||
|
||||
msg = OutboundMessage(
|
||||
channel=channel,
|
||||
chat_id=chat_id,
|
||||
content=content
|
||||
content=content,
|
||||
media=media or []
|
||||
)
|
||||
|
||||
try:
|
||||
|
||||
@@ -32,20 +32,33 @@ class ToolRegistry:
|
||||
return name in self._tools
|
||||
|
||||
def get_definitions(self) -> list[dict[str, Any]]:
|
||||
"""Get all tool definitions in OpenAI format."""
|
||||
return [tool.to_schema() for tool in self._tools.values()]
|
||||
"""Get tool definitions for all registered tools.
|
||||
|
||||
Supports both function tools (with to_schema) and native tools (with to_params).
|
||||
"""
|
||||
definitions = []
|
||||
for tool in self._tools.values():
|
||||
if hasattr(tool, 'to_params'): # Native Anthropic tool
|
||||
definitions.append(tool.to_params())
|
||||
elif hasattr(tool, 'to_schema'): # Function tool
|
||||
definitions.append(tool.to_schema())
|
||||
else:
|
||||
raise ValueError(f"Tool {tool.name} has no schema method (to_params or to_schema)")
|
||||
return definitions
|
||||
|
||||
async def execute(self, name: str, params: dict[str, Any]) -> str:
|
||||
async def execute(self, name: str, params: dict[str, Any]) -> Any:
|
||||
"""
|
||||
Execute a tool by name with given parameters.
|
||||
|
||||
|
||||
Supports both native Anthropic tools (via __call__) and function tools (via execute).
|
||||
|
||||
Args:
|
||||
name: Tool name.
|
||||
params: Tool parameters.
|
||||
|
||||
|
||||
Returns:
|
||||
Tool execution result as string.
|
||||
|
||||
Tool execution result (ToolResult, CLIResult, or string).
|
||||
|
||||
Raises:
|
||||
KeyError: If tool not found.
|
||||
"""
|
||||
@@ -54,20 +67,33 @@ class ToolRegistry:
|
||||
return f"Error: Tool '{name}' not found"
|
||||
|
||||
try:
|
||||
errors = tool.validate_params(params)
|
||||
if errors:
|
||||
return f"Error: Invalid parameters for tool '{name}': " + "; ".join(errors)
|
||||
return await tool.execute(**params)
|
||||
# Duck typing - support both native and function tools
|
||||
if hasattr(tool, 'to_params'):
|
||||
# Native Anthropic tool - call directly via __call__, no validation needed
|
||||
return await tool(**params)
|
||||
else:
|
||||
# Legacy function tool - validate then execute
|
||||
errors = tool.validate_params(params)
|
||||
if errors:
|
||||
return f"Error: Invalid parameters for tool '{name}': " + "; ".join(errors)
|
||||
return await tool.execute(**params)
|
||||
except Exception as e:
|
||||
return f"Error executing {name}: {str(e)}"
|
||||
|
||||
def get_tools(self) -> list[Any]:
|
||||
"""Get list of tool objects (not definitions).
|
||||
|
||||
Returns tool objects which can be inspected for metadata like beta_flag.
|
||||
"""
|
||||
return list(self._tools.values())
|
||||
|
||||
@property
|
||||
def tool_names(self) -> list[str]:
|
||||
"""Get list of registered tool names."""
|
||||
return list(self._tools.keys())
|
||||
|
||||
|
||||
def __len__(self) -> int:
|
||||
return len(self._tools)
|
||||
|
||||
|
||||
def __contains__(self, name: str) -> bool:
|
||||
return name in self._tools
|
||||
|
||||
@@ -20,11 +20,13 @@ class SpawnTool(Tool):
|
||||
self._manager = manager
|
||||
self._origin_channel = "cli"
|
||||
self._origin_chat_id = "direct"
|
||||
|
||||
def set_context(self, channel: str, chat_id: str) -> None:
|
||||
self._origin_metadata: dict[str, Any] = {}
|
||||
|
||||
def set_context(self, channel: str, chat_id: str, metadata: dict[str, Any] | None = None) -> None:
|
||||
"""Set the origin context for subagent announcements."""
|
||||
self._origin_channel = channel
|
||||
self._origin_chat_id = chat_id
|
||||
self._origin_metadata = metadata or {}
|
||||
|
||||
@property
|
||||
def name(self) -> str:
|
||||
@@ -67,4 +69,5 @@ class SpawnTool(Tool):
|
||||
model=model,
|
||||
origin_channel=self._origin_channel,
|
||||
origin_chat_id=self._origin_chat_id,
|
||||
origin_metadata=self._origin_metadata,
|
||||
)
|
||||
|
||||
@@ -0,0 +1,72 @@
|
||||
"""Message tool for subagents to communicate with the main agent."""
|
||||
|
||||
from typing import Any, TYPE_CHECKING
|
||||
|
||||
from nanobot.agent.tools.base import Tool
|
||||
from nanobot.bus.events import InboundMessage
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from nanobot.bus.queue import MessageBus
|
||||
|
||||
|
||||
class SubagentMessageTool(Tool):
|
||||
"""
|
||||
Tool for subagents to send messages to the main agent.
|
||||
|
||||
Messages are sent via the bus and preserve metadata (e.g. suppress_output)
|
||||
from the originating message that spawned the subagent.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
bus: "MessageBus",
|
||||
origin_channel: str,
|
||||
origin_chat_id: str,
|
||||
origin_metadata: dict[str, Any] | None = None,
|
||||
):
|
||||
self._bus = bus
|
||||
self._origin_channel = origin_channel
|
||||
self._origin_chat_id = origin_chat_id
|
||||
self._origin_metadata = origin_metadata or {}
|
||||
|
||||
@property
|
||||
def name(self) -> str:
|
||||
return "message"
|
||||
|
||||
@property
|
||||
def description(self) -> str:
|
||||
return (
|
||||
"Send a message to the main agent. "
|
||||
"Use this to communicate findings, request clarification, or provide updates. "
|
||||
"The main agent will process your message and decide how to respond."
|
||||
)
|
||||
|
||||
@property
|
||||
def parameters(self) -> dict[str, Any]:
|
||||
return {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"content": {
|
||||
"type": "string",
|
||||
"description": "The message content to send to the main agent"
|
||||
},
|
||||
},
|
||||
"required": ["content"]
|
||||
}
|
||||
|
||||
async def execute(self, content: str, **kwargs: Any) -> str:
|
||||
"""Send a message to the main agent via the bus."""
|
||||
# Create InboundMessage to trigger main agent
|
||||
msg = InboundMessage(
|
||||
channel="system",
|
||||
sender_id="subagent",
|
||||
chat_id=f"{self._origin_channel}:{self._origin_chat_id}",
|
||||
content=f"[Subagent message]\n\n{content}",
|
||||
metadata=self._origin_metadata,
|
||||
)
|
||||
|
||||
try:
|
||||
await self._bus.publish_inbound(msg)
|
||||
return "Message sent to main agent"
|
||||
except Exception as e:
|
||||
return f"Error sending message: {str(e)}"
|
||||
@@ -4,6 +4,8 @@ from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import re
|
||||
from pathlib import Path
|
||||
|
||||
from loguru import logger
|
||||
from telegram import BotCommand, Update
|
||||
from telegram.ext import Application, CommandHandler, MessageHandler, filters, ContextTypes
|
||||
@@ -199,11 +201,17 @@ class TelegramChannel(BaseChannel):
|
||||
chat_id = int(msg.chat_id)
|
||||
# Convert markdown to Telegram HTML
|
||||
html_content = _markdown_to_telegram_html(msg.content)
|
||||
await self._app.bot.send_message(
|
||||
chat_id=chat_id,
|
||||
text=html_content,
|
||||
parse_mode="HTML"
|
||||
)
|
||||
|
||||
# Check if message has media attachments
|
||||
if msg.media:
|
||||
await self._send_with_media(chat_id, html_content, msg.media)
|
||||
else:
|
||||
# Text-only message
|
||||
await self._app.bot.send_message(
|
||||
chat_id=chat_id,
|
||||
text=html_content,
|
||||
parse_mode="HTML"
|
||||
)
|
||||
except ValueError:
|
||||
logger.error(f"Invalid chat_id: {msg.chat_id}")
|
||||
except Exception as e:
|
||||
@@ -216,7 +224,167 @@ class TelegramChannel(BaseChannel):
|
||||
)
|
||||
except Exception as e2:
|
||||
logger.error(f"Error sending Telegram message: {e2}")
|
||||
|
||||
|
||||
async def _send_with_media(self, chat_id: int, caption: str, media_paths: list[str]) -> None:
|
||||
"""
|
||||
Send message with media attachments.
|
||||
|
||||
Args:
|
||||
chat_id: Telegram chat ID
|
||||
caption: Message caption
|
||||
media_paths: List of file paths or URLs
|
||||
"""
|
||||
from telegram import InputMediaPhoto, InputMediaVideo
|
||||
|
||||
from nanobot.channels.telegram_media import (
|
||||
MediaKind,
|
||||
classify_media,
|
||||
detect_mime,
|
||||
fetch_media,
|
||||
group_media_for_album,
|
||||
optimize_image,
|
||||
)
|
||||
|
||||
# Process each media item
|
||||
processed_media: list[tuple[str, MediaKind, bytes, str]] = []
|
||||
|
||||
for path in media_paths:
|
||||
try:
|
||||
# Fetch remote URLs
|
||||
if path.startswith(("http://", "https://")):
|
||||
content, mime = await fetch_media(path, max_bytes=100_000_000)
|
||||
kind = classify_media(mime)
|
||||
# Extract filename from URL
|
||||
filename = Path(path).name
|
||||
else:
|
||||
# Local file
|
||||
file_path = Path(path)
|
||||
if not file_path.exists():
|
||||
logger.warning(f"Media file not found: {path}")
|
||||
continue
|
||||
|
||||
with open(file_path, "rb") as f:
|
||||
content = f.read()
|
||||
|
||||
mime = detect_mime(path, content)
|
||||
kind = classify_media(mime)
|
||||
# Extract filename from local path
|
||||
filename = file_path.name
|
||||
|
||||
# Optimize images
|
||||
if kind == MediaKind.IMAGE:
|
||||
try:
|
||||
content = optimize_image(path, max_bytes=6_000_000)
|
||||
except Exception as e:
|
||||
logger.warning(f"Image optimization failed: {e}, sending original")
|
||||
|
||||
processed_media.append((path, kind, content, filename))
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to process media {path}: {e}")
|
||||
continue
|
||||
|
||||
if not processed_media:
|
||||
# No media could be processed, send text only
|
||||
await self._app.bot.send_message(
|
||||
chat_id=chat_id,
|
||||
text=caption,
|
||||
parse_mode="HTML"
|
||||
)
|
||||
return
|
||||
|
||||
# Group media for album sending
|
||||
media_items = [(path, kind) for path, kind, _, _ in processed_media]
|
||||
grouping = group_media_for_album(media_items)
|
||||
|
||||
# Handle caption length (Telegram limit: 1024 chars)
|
||||
if len(caption) > 1024:
|
||||
# Send media without caption, then follow-up text
|
||||
media_caption = None
|
||||
followup_text = caption
|
||||
else:
|
||||
media_caption = caption
|
||||
followup_text = None
|
||||
|
||||
# Send album if grouped
|
||||
if grouping["album"]:
|
||||
album_paths = grouping["album"]
|
||||
album_media = []
|
||||
|
||||
for path, kind, content, filename in processed_media:
|
||||
if path not in album_paths:
|
||||
continue
|
||||
|
||||
if kind == MediaKind.IMAGE:
|
||||
media_obj = InputMediaPhoto(
|
||||
media=content,
|
||||
caption=media_caption if len(album_media) == 0 else None,
|
||||
parse_mode="HTML" if media_caption else None
|
||||
)
|
||||
elif kind == MediaKind.VIDEO:
|
||||
media_obj = InputMediaVideo(
|
||||
media=content,
|
||||
caption=media_caption if len(album_media) == 0 else None,
|
||||
parse_mode="HTML" if media_caption else None
|
||||
)
|
||||
else:
|
||||
continue # Skip non-album types
|
||||
|
||||
album_media.append(media_obj)
|
||||
|
||||
if album_media:
|
||||
await self._app.bot.send_media_group(
|
||||
chat_id=chat_id,
|
||||
media=album_media
|
||||
)
|
||||
|
||||
# Send separate media
|
||||
for i, (path, kind, content, filename) in enumerate(processed_media):
|
||||
if path in grouping["album"]:
|
||||
continue # Already sent in album
|
||||
|
||||
# Only first separate item gets caption
|
||||
item_caption = media_caption if i == 0 else None
|
||||
|
||||
if kind == MediaKind.IMAGE:
|
||||
await self._app.bot.send_photo(
|
||||
chat_id=chat_id,
|
||||
photo=content,
|
||||
caption=item_caption,
|
||||
parse_mode="HTML" if item_caption else None
|
||||
)
|
||||
elif kind == MediaKind.VIDEO:
|
||||
await self._app.bot.send_video(
|
||||
chat_id=chat_id,
|
||||
video=content,
|
||||
caption=item_caption,
|
||||
parse_mode="HTML" if item_caption else None
|
||||
)
|
||||
elif kind == MediaKind.AUDIO:
|
||||
await self._app.bot.send_audio(
|
||||
chat_id=chat_id,
|
||||
audio=content,
|
||||
caption=item_caption,
|
||||
parse_mode="HTML" if item_caption else None,
|
||||
filename=filename
|
||||
)
|
||||
elif kind == MediaKind.DOCUMENT:
|
||||
await self._app.bot.send_document(
|
||||
chat_id=chat_id,
|
||||
document=content,
|
||||
caption=item_caption,
|
||||
parse_mode="HTML" if item_caption else None,
|
||||
filename=filename
|
||||
)
|
||||
|
||||
# Send follow-up text if caption was too long
|
||||
if followup_text:
|
||||
await self._app.bot.send_message(
|
||||
chat_id=chat_id,
|
||||
text=followup_text,
|
||||
parse_mode="HTML"
|
||||
)
|
||||
|
||||
async def _on_start(self, update: Update, context: ContextTypes.DEFAULT_TYPE) -> None:
|
||||
"""Handle /start command."""
|
||||
if not update.message or not update.effective_user:
|
||||
|
||||
@@ -0,0 +1,286 @@
|
||||
"""Media handling utilities for Telegram channel."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import io
|
||||
import mimetypes
|
||||
from enum import Enum
|
||||
from pathlib import Path
|
||||
|
||||
import httpx
|
||||
from loguru import logger
|
||||
from PIL import Image
|
||||
|
||||
# Telegram API photo size limit (6MB)
|
||||
TELEGRAM_PHOTO_SIZE_LIMIT = 6_000_000
|
||||
|
||||
try:
|
||||
import magic
|
||||
HAS_MAGIC = True
|
||||
except ImportError:
|
||||
HAS_MAGIC = False
|
||||
|
||||
try:
|
||||
from pillow_heif import register_heif_opener
|
||||
register_heif_opener()
|
||||
HAS_HEIF = True
|
||||
except ImportError:
|
||||
HAS_HEIF = False
|
||||
|
||||
|
||||
class MediaKind(Enum):
|
||||
"""Media type classification."""
|
||||
IMAGE = "image"
|
||||
VIDEO = "video"
|
||||
AUDIO = "audio"
|
||||
DOCUMENT = "document"
|
||||
|
||||
|
||||
def detect_mime(path: str, content: bytes | None = None) -> str:
|
||||
"""
|
||||
Detect MIME type of media file.
|
||||
|
||||
Priority:
|
||||
1. python-magic sniff (if available and content provided)
|
||||
2. Extension-based lookup
|
||||
3. Fallback to application/octet-stream
|
||||
|
||||
Args:
|
||||
path: File path (used for extension detection)
|
||||
content: Optional file content bytes for magic sniffing
|
||||
|
||||
Returns:
|
||||
MIME type string (e.g., "image/jpeg")
|
||||
"""
|
||||
# Try magic detection first if we have content
|
||||
if HAS_MAGIC and content:
|
||||
try:
|
||||
mime = magic.from_buffer(content, mime=True)
|
||||
# Avoid generic types if we can be more specific from extension
|
||||
if mime and mime != "application/octet-stream":
|
||||
return mime
|
||||
except Exception as e:
|
||||
logger.debug(f"Magic detection failed, falling back to extension: {e}")
|
||||
|
||||
# Extension-based detection
|
||||
mime_type, _ = mimetypes.guess_type(path)
|
||||
if mime_type:
|
||||
return mime_type
|
||||
|
||||
# Fallback
|
||||
return "application/octet-stream"
|
||||
|
||||
|
||||
def classify_media(mime: str) -> MediaKind:
|
||||
"""
|
||||
Classify MIME type into media kind.
|
||||
|
||||
Args:
|
||||
mime: MIME type string (e.g., "image/jpeg")
|
||||
|
||||
Returns:
|
||||
MediaKind enum value
|
||||
"""
|
||||
if mime.startswith("image/"):
|
||||
return MediaKind.IMAGE
|
||||
if mime.startswith("video/"):
|
||||
return MediaKind.VIDEO
|
||||
if mime.startswith("audio/"):
|
||||
return MediaKind.AUDIO
|
||||
# Everything else is a document
|
||||
return MediaKind.DOCUMENT
|
||||
|
||||
|
||||
def is_heic_format(path: str) -> bool:
|
||||
"""
|
||||
Check if file is HEIC/HEIF format.
|
||||
|
||||
Args:
|
||||
path: File path
|
||||
|
||||
Returns:
|
||||
True if file extension is .heic or .heif
|
||||
"""
|
||||
ext = Path(path).suffix.lower()
|
||||
return ext in (".heic", ".heif")
|
||||
|
||||
|
||||
async def fetch_media(url: str, max_bytes: int) -> tuple[bytes, str]:
|
||||
"""
|
||||
Download media from remote URL.
|
||||
|
||||
Args:
|
||||
url: Remote URL to fetch
|
||||
max_bytes: Maximum size to download
|
||||
|
||||
Returns:
|
||||
Tuple of (content bytes, detected MIME type)
|
||||
|
||||
Raises:
|
||||
ValueError: If download fails or exceeds size limit
|
||||
"""
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=10.0) as client:
|
||||
response = await client.get(url, follow_redirects=True)
|
||||
response.raise_for_status()
|
||||
|
||||
content = response.content
|
||||
|
||||
if len(content) > max_bytes:
|
||||
raise ValueError(f"Media exceeds size limit: {len(content)} > {max_bytes}")
|
||||
|
||||
# Get MIME type from response or detect
|
||||
mime = response.headers.get("content-type", "application/octet-stream")
|
||||
# Strip charset if present (e.g., "image/jpeg; charset=utf-8" → "image/jpeg")
|
||||
mime = mime.split(";")[0].strip()
|
||||
|
||||
# Detect from content if generic type
|
||||
if mime == "application/octet-stream":
|
||||
mime = detect_mime(url, content)
|
||||
|
||||
return content, mime
|
||||
|
||||
except httpx.TimeoutException as e:
|
||||
raise ValueError(f"Download timeout: {url}") from e
|
||||
except httpx.HTTPError as e:
|
||||
raise ValueError(f"Download failed: {url}: {e}") from e
|
||||
|
||||
|
||||
def optimize_image(path: str, max_bytes: int = TELEGRAM_PHOTO_SIZE_LIMIT) -> bytes:
|
||||
"""
|
||||
Optimize image to fit under size limit.
|
||||
|
||||
Strategy:
|
||||
1. Convert HEIC to JPEG if needed
|
||||
2. PNG with alpha → preserve with compression levels [6,7,8,9]
|
||||
3. JPEG/PNG without alpha → resize + quality grid
|
||||
|
||||
Sizes: [2048, 1536, 1280, 1024, 800] px (max dimension)
|
||||
Qualities: [80, 70, 60, 50, 40] (JPEG only)
|
||||
|
||||
Args:
|
||||
path: Path to image file
|
||||
max_bytes: Maximum size in bytes (default 6MB for Telegram)
|
||||
|
||||
Returns:
|
||||
Optimized image bytes
|
||||
|
||||
Raises:
|
||||
ValueError: If image cannot be optimized under limit
|
||||
"""
|
||||
# Load image with context manager to ensure file handle is closed
|
||||
with Image.open(path) as img:
|
||||
# Convert HEIC to JPEG
|
||||
if is_heic_format(path):
|
||||
if not HAS_HEIF:
|
||||
raise ValueError("pillow-heif not available for HEIC conversion")
|
||||
# Convert to RGB (HEIC → JPEG)
|
||||
if img.mode != "RGB":
|
||||
img = img.convert("RGB")
|
||||
return _optimize_jpeg(img, max_bytes)
|
||||
|
||||
# PNG with alpha channel - preserve it
|
||||
if img.mode == "RGBA" or img.mode == "LA":
|
||||
return _optimize_png(img, max_bytes)
|
||||
|
||||
# Everything else → convert to JPEG and optimize
|
||||
if img.mode != "RGB":
|
||||
img = img.convert("RGB")
|
||||
return _optimize_jpeg(img, max_bytes)
|
||||
|
||||
|
||||
def _optimize_jpeg(img: Image.Image, max_bytes: int) -> bytes:
|
||||
"""Optimize JPEG with size/quality grid."""
|
||||
sizes = [2048, 1536, 1280, 1024, 800]
|
||||
qualities = [80, 70, 60, 50, 40]
|
||||
|
||||
for size in sizes:
|
||||
# Always copy to avoid mutation issues
|
||||
resized = img.copy()
|
||||
if max(img.size) > size:
|
||||
resized.thumbnail((size, size), Image.Resampling.LANCZOS)
|
||||
|
||||
for quality in qualities:
|
||||
buf = io.BytesIO()
|
||||
resized.save(buf, format="JPEG", quality=quality, optimize=True)
|
||||
data = buf.getvalue()
|
||||
|
||||
if len(data) <= max_bytes:
|
||||
return data
|
||||
|
||||
# If we get here, even smallest size/quality is too large
|
||||
raise ValueError(f"Cannot optimize image under {max_bytes} bytes")
|
||||
|
||||
|
||||
def _optimize_png(img: Image.Image, max_bytes: int) -> bytes:
|
||||
"""Optimize PNG while preserving alpha channel."""
|
||||
compress_levels = [6, 7, 8, 9]
|
||||
sizes = [2048, 1536, 1280, 1024, 800]
|
||||
|
||||
for size in sizes:
|
||||
# Always copy to avoid mutation issues
|
||||
resized = img.copy()
|
||||
if max(img.size) > size:
|
||||
resized.thumbnail((size, size), Image.Resampling.LANCZOS)
|
||||
|
||||
for compress_level in compress_levels:
|
||||
buf = io.BytesIO()
|
||||
resized.save(buf, format="PNG", compress_level=compress_level, optimize=True)
|
||||
data = buf.getvalue()
|
||||
|
||||
if len(data) <= max_bytes:
|
||||
return data
|
||||
|
||||
# Fallback: try converting to JPEG if still too large
|
||||
if img.mode in ("RGBA", "LA"):
|
||||
# Create white background
|
||||
background = Image.new("RGB", img.size, (255, 255, 255))
|
||||
if img.mode == "RGBA":
|
||||
background.paste(img, mask=img.split()[3]) # Use alpha as mask
|
||||
else: # LA (grayscale + alpha)
|
||||
background.paste(img.convert("L"), mask=img.split()[1])
|
||||
return _optimize_jpeg(background, max_bytes)
|
||||
|
||||
raise ValueError(f"Cannot optimize PNG under {max_bytes} bytes")
|
||||
|
||||
|
||||
def group_media_for_album(media_items: list[tuple[str, MediaKind]]) -> dict[str, list[str]]:
|
||||
"""
|
||||
Group media items for album sending.
|
||||
|
||||
Logic:
|
||||
- All images (2+) → album
|
||||
- All videos (2+) → album
|
||||
- Mixed types → separate
|
||||
- Single item → separate
|
||||
|
||||
Args:
|
||||
media_items: List of (path, MediaKind) tuples
|
||||
|
||||
Returns:
|
||||
Dict with 'album' and 'separate' keys containing lists of paths
|
||||
"""
|
||||
if len(media_items) <= 1:
|
||||
return {
|
||||
"album": [],
|
||||
"separate": [path for path, _ in media_items]
|
||||
}
|
||||
|
||||
# Count each kind
|
||||
kinds = [kind for _, kind in media_items]
|
||||
unique_kinds = set(kinds)
|
||||
|
||||
# All same type → album (if images or videos)
|
||||
if len(unique_kinds) == 1:
|
||||
kind = kinds[0]
|
||||
if kind in (MediaKind.IMAGE, MediaKind.VIDEO):
|
||||
return {
|
||||
"album": [path for path, _ in media_items],
|
||||
"separate": []
|
||||
}
|
||||
|
||||
# Mixed types or non-album-able types → separate
|
||||
return {
|
||||
"album": [],
|
||||
"separate": [path for path, _ in media_items]
|
||||
}
|
||||
+24
-1
@@ -4,6 +4,7 @@ import asyncio
|
||||
import os
|
||||
import select
|
||||
import signal
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
@@ -294,6 +295,26 @@ def _make_provider(config):
|
||||
# ============================================================================
|
||||
|
||||
|
||||
def _start_moltbook_loop():
|
||||
"""Start the moltbook polling loop in the background."""
|
||||
loop_script = Path.home() / ".nanobot" / "scripts" / "moltbook-loop.sh"
|
||||
log_file = Path.home() / ".nanobot" / "scripts" / "moltbook-loop.log"
|
||||
|
||||
if not loop_script.exists():
|
||||
return
|
||||
|
||||
try:
|
||||
subprocess.Popen(
|
||||
["/bin/bash", str(loop_script)],
|
||||
stdout=open(log_file, "a"),
|
||||
stderr=subprocess.STDOUT,
|
||||
start_new_session=True,
|
||||
)
|
||||
console.print(f"[green]✓[/green] Moltbook polling: every 15m")
|
||||
except Exception as e:
|
||||
console.print(f"[yellow]Warning: Could not start moltbook loop: {e}[/yellow]")
|
||||
|
||||
|
||||
@app.command()
|
||||
def gateway(
|
||||
port: int = typer.Option(18790, "--port", "-p", help="Gateway port"),
|
||||
@@ -376,7 +397,7 @@ def gateway(
|
||||
enabled=True,
|
||||
session_manager=session_manager, # Pass session manager
|
||||
target_session_key="telegram:239824268", # Target session
|
||||
idle_threshold_s=30 * 60, # 30 minutes idle
|
||||
idle_threshold_s=20 * 60, # 20 minutes idle
|
||||
)
|
||||
|
||||
# Create channel manager
|
||||
@@ -414,6 +435,8 @@ def gateway(
|
||||
|
||||
console.print("[green]✓[/green] Heartbeat: every 30m")
|
||||
|
||||
_start_moltbook_loop()
|
||||
|
||||
async def run():
|
||||
try:
|
||||
await cron.start()
|
||||
|
||||
@@ -118,10 +118,18 @@ class HeartbeatService:
|
||||
try:
|
||||
session = self.session_manager.get_or_create(self.target_session_key)
|
||||
|
||||
# Find last user message timestamp
|
||||
# Find last real user message timestamp (exclude system-generated messages)
|
||||
# Real Telegram messages have sender_id like "239824268|username"
|
||||
# System messages (heartbeat, cron) created via process_direct have sender_id="user"
|
||||
# Old messages may not have sender_id field (backwards compat: treat as real user messages)
|
||||
last_user_timestamp = None
|
||||
for msg in reversed(session.messages):
|
||||
if msg.get("role") == "user":
|
||||
sender_id = msg.get("sender_id")
|
||||
# Skip if explicitly marked as system-generated
|
||||
if sender_id == "user":
|
||||
continue
|
||||
# Accept if no sender_id (old message) or if real user ID
|
||||
last_user_timestamp = msg.get("timestamp")
|
||||
break
|
||||
|
||||
|
||||
@@ -199,21 +199,41 @@ class AnthropicOAuthProvider(LLMProvider):
|
||||
|
||||
def _convert_tools_to_anthropic(
|
||||
self,
|
||||
tools: list[dict[str, Any]] | None
|
||||
tools: list[dict[str, Any]] | list[Any] | None
|
||||
) -> list[dict[str, Any]] | None:
|
||||
"""Convert OpenAI-format tools to Anthropic format."""
|
||||
"""Convert tools to Anthropic API format.
|
||||
|
||||
Supports both function tools (custom) and native tools (Anthropic).
|
||||
Function tools are converted to Anthropic format.
|
||||
Native tools are passed through unchanged.
|
||||
Tool objects (with to_params/to_schema methods) are converted to dicts.
|
||||
"""
|
||||
if not tools:
|
||||
return None
|
||||
|
||||
anthropic_tools = []
|
||||
for tool in tools:
|
||||
if tool.get("type") == "function":
|
||||
func = tool["function"]
|
||||
# Convert tool objects to dicts first
|
||||
if hasattr(tool, 'to_params'): # Native Anthropic tool
|
||||
tool_dict = tool.to_params()
|
||||
elif hasattr(tool, 'to_schema'): # Function tool
|
||||
tool_dict = tool.to_schema()
|
||||
else:
|
||||
tool_dict = tool # Already a dict
|
||||
|
||||
# Now process the dict
|
||||
if tool_dict.get("type") == "function":
|
||||
# Convert function tool format
|
||||
func = tool_dict["function"]
|
||||
anthropic_tools.append({
|
||||
"name": func["name"],
|
||||
"description": func.get("description", ""),
|
||||
"input_schema": func.get("parameters", {"type": "object", "properties": {}})
|
||||
})
|
||||
else:
|
||||
# Pass through native tool format as-is
|
||||
# (bash_20250124, text_editor_20250728, computer_20251124, etc.)
|
||||
anthropic_tools.append(tool_dict)
|
||||
|
||||
return anthropic_tools if anthropic_tools else None
|
||||
|
||||
@@ -227,6 +247,7 @@ class AnthropicOAuthProvider(LLMProvider):
|
||||
tools: list[dict[str, Any]] | None = None,
|
||||
thinking_budget_override: int | None = None,
|
||||
context_management: dict[str, Any] | None = None,
|
||||
beta_flags: set[str] | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""Make request to Anthropic API."""
|
||||
client = await self._get_client()
|
||||
@@ -276,17 +297,28 @@ class AnthropicOAuthProvider(LLMProvider):
|
||||
payload["context_management"] = context_management
|
||||
|
||||
edit_types = [e.get("type") for e in (context_management or {}).get("edits", [])]
|
||||
|
||||
# Build headers with beta flags if provided
|
||||
headers = self._get_headers()
|
||||
if beta_flags:
|
||||
# Merge with existing beta header (from OAuth hardcoded flags)
|
||||
existing_beta = headers.get("anthropic-beta", "")
|
||||
existing_flags = set(existing_beta.split(",")) if existing_beta else set()
|
||||
all_flags = existing_flags | beta_flags
|
||||
headers["anthropic-beta"] = ",".join(sorted(all_flags))
|
||||
|
||||
logger.info(
|
||||
"Anthropic request: model={} max_tokens={} thinking={} tools={} context_mgmt={}",
|
||||
"Anthropic request: model={} max_tokens={} thinking={} tools={} context_mgmt={} beta={}",
|
||||
payload.get("model"), payload.get("max_tokens"),
|
||||
payload.get("thinking", "disabled"),
|
||||
len(payload.get("tools", [])),
|
||||
edit_types or "none",
|
||||
headers.get("anthropic-beta", "none"),
|
||||
)
|
||||
|
||||
response = await client.post(
|
||||
self._get_api_url(),
|
||||
headers=self._get_headers(),
|
||||
headers=headers,
|
||||
json=payload,
|
||||
)
|
||||
|
||||
@@ -332,7 +364,7 @@ class AnthropicOAuthProvider(LLMProvider):
|
||||
async def chat(
|
||||
self,
|
||||
messages: list[dict[str, Any]],
|
||||
tools: list[dict[str, Any]] | None = None,
|
||||
tools: list[dict[str, Any]] | list[Any] | None = None,
|
||||
model: str | None = None,
|
||||
max_tokens: int = 4096,
|
||||
temperature: float = 0.7,
|
||||
@@ -350,6 +382,17 @@ class AnthropicOAuthProvider(LLMProvider):
|
||||
model = self._normalize_model(model)
|
||||
|
||||
system, prepared_messages = self._prepare_messages(messages)
|
||||
|
||||
# Collect beta flags from native tools BEFORE conversion
|
||||
beta_flags: set[str] = set()
|
||||
if tools:
|
||||
for tool in tools:
|
||||
if hasattr(tool, 'beta_flag') and tool.beta_flag:
|
||||
beta_flags.add(tool.beta_flag)
|
||||
|
||||
logger.debug(f"Beta flags collected: {beta_flags} (from {len(tools) if tools else 0} tools)")
|
||||
|
||||
# Convert tools to API format
|
||||
anthropic_tools = self._convert_tools_to_anthropic(tools)
|
||||
|
||||
# Per-call thinking override (None = use instance default)
|
||||
@@ -365,6 +408,7 @@ class AnthropicOAuthProvider(LLMProvider):
|
||||
tools=anthropic_tools,
|
||||
thinking_budget_override=effective_thinking,
|
||||
context_management=context_management,
|
||||
beta_flags=beta_flags,
|
||||
)
|
||||
return self._parse_response(response)
|
||||
except Exception as e:
|
||||
|
||||
@@ -38,6 +38,7 @@ dependencies = [
|
||||
"qq-botpy>=1.0.0",
|
||||
"python-socks[asyncio]>=2.4.0",
|
||||
"prompt-toolkit>=3.0.0",
|
||||
"vncdotool>=1.0.0",
|
||||
]
|
||||
|
||||
[project.optional-dependencies]
|
||||
|
||||
@@ -45,7 +45,7 @@ async def test_process_direct_passes_metadata():
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_suppress_mode_adds_hidden_prefix():
|
||||
"""Test that suppress_output metadata adds [HIDDEN] prefix."""
|
||||
"""Test that suppress_output metadata adds [HIDDEN:signature] prefix."""
|
||||
bus = MessageBus()
|
||||
provider = MagicMock(spec=LLMProvider)
|
||||
provider.chat = AsyncMock(return_value=LLMResponse(
|
||||
@@ -66,14 +66,20 @@ async def test_suppress_mode_adds_hidden_prefix():
|
||||
metadata={"suppress_output": True}
|
||||
)
|
||||
|
||||
# Response content should have [HIDDEN] prefix
|
||||
assert response.startswith("[HIDDEN]")
|
||||
# Response content should have [HIDDEN:signature] prefix with 8-char hex signature
|
||||
assert response.startswith("[HIDDEN:")
|
||||
assert "]" in response
|
||||
# Extract signature part between [HIDDEN: and ]
|
||||
prefix_end = response.index("]")
|
||||
signature = response[8:prefix_end] # Skip "[HIDDEN:" to get signature
|
||||
assert len(signature) == 8 # 8-character hex signature
|
||||
assert all(c in "0123456789abcdef" for c in signature) # Valid hex
|
||||
assert "This is the agent response" in response
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_normal_mode_no_hidden_prefix():
|
||||
"""Test that normal messages don't get [HIDDEN] prefix."""
|
||||
"""Test that normal messages don't get [HIDDEN:signature] prefix."""
|
||||
bus = MessageBus()
|
||||
provider = MagicMock(spec=LLMProvider)
|
||||
provider.chat = AsyncMock(return_value=LLMResponse(
|
||||
@@ -91,6 +97,6 @@ async def test_normal_mode_no_hidden_prefix():
|
||||
# Call without suppress_output
|
||||
response = await loop.process_direct(content="test message")
|
||||
|
||||
# Response should NOT have [HIDDEN] prefix
|
||||
assert not response.startswith("[HIDDEN]")
|
||||
# Response should NOT have [HIDDEN:signature] prefix
|
||||
assert not response.startswith("[HIDDEN:")
|
||||
assert response == "Normal response"
|
||||
|
||||
@@ -0,0 +1,262 @@
|
||||
"""Tests for agent loop handling of ToolResult and CLIResult objects."""
|
||||
|
||||
import pytest
|
||||
from pathlib import Path
|
||||
from unittest.mock import AsyncMock, MagicMock, patch
|
||||
|
||||
from nanobot.agent.loop import AgentLoop
|
||||
from nanobot.agent.tools.anthropic.base import ToolResult, CLIResult
|
||||
from nanobot.bus.events import InboundMessage, OutboundMessage
|
||||
from nanobot.bus.queue import MessageBus
|
||||
from nanobot.providers.base import LLMResponse, ToolCallRequest
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def mock_provider():
|
||||
"""Create mock LLM provider."""
|
||||
provider = MagicMock()
|
||||
provider.chat = AsyncMock()
|
||||
provider.thinking_budget = 0
|
||||
return provider
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def mock_session_manager():
|
||||
"""Create mock session manager."""
|
||||
session_mgr = MagicMock()
|
||||
session_mgr.load = AsyncMock(return_value={
|
||||
"messages": [],
|
||||
"metadata": {},
|
||||
})
|
||||
session_mgr.save = AsyncMock()
|
||||
return session_mgr
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def mock_bus():
|
||||
"""Create mock message bus."""
|
||||
bus = MagicMock(spec=MessageBus)
|
||||
bus.publish = AsyncMock()
|
||||
return bus
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def agent_loop(mock_provider, mock_session_manager, mock_bus, tmp_path):
|
||||
"""Create agent loop for testing."""
|
||||
return AgentLoop(
|
||||
provider=mock_provider,
|
||||
session_manager=mock_session_manager,
|
||||
bus=mock_bus,
|
||||
workspace=tmp_path,
|
||||
max_iterations=5,
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_tool_result_with_output(agent_loop, mock_provider):
|
||||
"""Test handling ToolResult with output field."""
|
||||
# Mock LLM responses
|
||||
mock_provider.chat.side_effect = [
|
||||
# First call: request tool
|
||||
LLMResponse(
|
||||
content="Using tool",
|
||||
tool_calls=[ToolCallRequest(id="call_1", name="test_tool", arguments={})],
|
||||
),
|
||||
# Second call: final response
|
||||
LLMResponse(content="Done"),
|
||||
]
|
||||
|
||||
# Mock tool that returns ToolResult
|
||||
tool_result = ToolResult(output="Tool executed successfully")
|
||||
agent_loop.tools.execute = AsyncMock(return_value=tool_result)
|
||||
|
||||
message = InboundMessage(
|
||||
channel="test",
|
||||
chat_id="123",
|
||||
sender_id="user1",
|
||||
content="Test message",
|
||||
)
|
||||
|
||||
response = await agent_loop._process_message(message)
|
||||
|
||||
# Verify tool result was added to messages
|
||||
calls = mock_provider.chat.call_args_list
|
||||
second_call_messages = calls[1][1]["messages"]
|
||||
|
||||
# Find the tool result message
|
||||
tool_msg = next(m for m in second_call_messages if m.get("role") == "tool")
|
||||
assert tool_msg["content"] == "Tool executed successfully"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_tool_result_with_error(agent_loop, mock_provider):
|
||||
"""Test handling ToolResult with error field."""
|
||||
mock_provider.chat.side_effect = [
|
||||
LLMResponse(
|
||||
content="Using tool",
|
||||
tool_calls=[ToolCallRequest(id="call_1", name="test_tool", arguments={})],
|
||||
),
|
||||
LLMResponse(content="Error handled"),
|
||||
]
|
||||
|
||||
tool_result = ToolResult(error="Command failed: exit code 1")
|
||||
agent_loop.tools.execute = AsyncMock(return_value=tool_result)
|
||||
|
||||
message = InboundMessage(
|
||||
channel="test",
|
||||
chat_id="123",
|
||||
sender_id="user1",
|
||||
content="Test message",
|
||||
)
|
||||
|
||||
response = await agent_loop._process_message(message)
|
||||
|
||||
calls = mock_provider.chat.call_args_list
|
||||
second_call_messages = calls[1][1]["messages"]
|
||||
tool_msg = next(m for m in second_call_messages if m.get("role") == "tool")
|
||||
assert "Error:" in tool_msg["content"]
|
||||
assert "Command failed: exit code 1" in tool_msg["content"]
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_tool_result_with_base64_image(agent_loop, mock_provider):
|
||||
"""Test handling ToolResult with base64_image field."""
|
||||
mock_provider.chat.side_effect = [
|
||||
LLMResponse(
|
||||
content="Taking screenshot",
|
||||
tool_calls=[ToolCallRequest(id="call_1", name="screenshot", arguments={})],
|
||||
),
|
||||
LLMResponse(content="Screenshot analyzed"),
|
||||
]
|
||||
|
||||
tool_result = ToolResult(
|
||||
output="Screenshot taken",
|
||||
base64_image="iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAYAAAAfFcSJAAAADUlEQVR42mNk+M9QDwADhgGAWjR9awAAAABJRU5ErkJggg==",
|
||||
)
|
||||
agent_loop.tools.execute = AsyncMock(return_value=tool_result)
|
||||
|
||||
message = InboundMessage(
|
||||
channel="test",
|
||||
chat_id="123",
|
||||
sender_id="user1",
|
||||
content="Test message",
|
||||
)
|
||||
|
||||
response = await agent_loop._process_message(message)
|
||||
|
||||
calls = mock_provider.chat.call_args_list
|
||||
second_call_messages = calls[1][1]["messages"]
|
||||
tool_msg = next(m for m in second_call_messages if m.get("role") == "tool")
|
||||
|
||||
# Should contain both text and image
|
||||
assert isinstance(tool_msg["content"], list)
|
||||
assert len(tool_msg["content"]) == 2
|
||||
|
||||
# Text content
|
||||
text_part = next(p for p in tool_msg["content"] if p["type"] == "text")
|
||||
assert text_part["text"] == "Screenshot taken"
|
||||
|
||||
# Image content
|
||||
image_part = next(p for p in tool_msg["content"] if p["type"] == "image")
|
||||
assert image_part["source"]["type"] == "base64"
|
||||
assert image_part["source"]["media_type"] == "image/png"
|
||||
assert "iVBORw0KGgoAAAANS" in image_part["source"]["data"]
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_cli_result_handling(agent_loop, mock_provider):
|
||||
"""Test handling CLIResult from text editor tools."""
|
||||
mock_provider.chat.side_effect = [
|
||||
LLMResponse(
|
||||
content="Editing file",
|
||||
tool_calls=[ToolCallRequest(id="call_1", name="edit", arguments={})],
|
||||
),
|
||||
LLMResponse(content="File edited"),
|
||||
]
|
||||
|
||||
cli_result = CLIResult(
|
||||
exit_code=0,
|
||||
output="File updated successfully",
|
||||
error="",
|
||||
)
|
||||
agent_loop.tools.execute = AsyncMock(return_value=cli_result)
|
||||
|
||||
message = InboundMessage(
|
||||
channel="test",
|
||||
chat_id="123",
|
||||
sender_id="user1",
|
||||
content="Test message",
|
||||
)
|
||||
|
||||
response = await agent_loop._process_message(message)
|
||||
|
||||
calls = mock_provider.chat.call_args_list
|
||||
second_call_messages = calls[1][1]["messages"]
|
||||
tool_msg = next(m for m in second_call_messages if m.get("role") == "tool")
|
||||
assert tool_msg["content"] == "File updated successfully"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_legacy_string_result(agent_loop, mock_provider):
|
||||
"""Test backward compatibility with string results from function tools."""
|
||||
mock_provider.chat.side_effect = [
|
||||
LLMResponse(
|
||||
content="Using tool",
|
||||
tool_calls=[ToolCallRequest(id="call_1", name="legacy_tool", arguments={})],
|
||||
),
|
||||
LLMResponse(content="Done"),
|
||||
]
|
||||
|
||||
# Legacy tool returns plain string
|
||||
agent_loop.tools.execute = AsyncMock(return_value="Plain text result")
|
||||
|
||||
message = InboundMessage(
|
||||
channel="test",
|
||||
chat_id="123",
|
||||
sender_id="user1",
|
||||
content="Test message",
|
||||
)
|
||||
|
||||
response = await agent_loop._process_message(message)
|
||||
|
||||
calls = mock_provider.chat.call_args_list
|
||||
second_call_messages = calls[1][1]["messages"]
|
||||
tool_msg = next(m for m in second_call_messages if m.get("role") == "tool")
|
||||
assert tool_msg["content"] == "Plain text result"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_tool_result_output_and_error(agent_loop, mock_provider):
|
||||
"""Test handling ToolResult with both output and error."""
|
||||
mock_provider.chat.side_effect = [
|
||||
LLMResponse(
|
||||
content="Running command",
|
||||
tool_calls=[ToolCallRequest(id="call_1", name="bash", arguments={})],
|
||||
),
|
||||
LLMResponse(content="Handled"),
|
||||
]
|
||||
|
||||
tool_result = ToolResult(
|
||||
output="Partial output before error",
|
||||
error="Unexpected termination",
|
||||
)
|
||||
agent_loop.tools.execute = AsyncMock(return_value=tool_result)
|
||||
|
||||
message = InboundMessage(
|
||||
channel="test",
|
||||
chat_id="123",
|
||||
sender_id="user1",
|
||||
content="Test message",
|
||||
)
|
||||
|
||||
response = await agent_loop._process_message(message)
|
||||
|
||||
calls = mock_provider.chat.call_args_list
|
||||
second_call_messages = calls[1][1]["messages"]
|
||||
tool_msg = next(m for m in second_call_messages if m.get("role") == "tool")
|
||||
|
||||
# Should contain both output and error
|
||||
content = tool_msg["content"]
|
||||
assert "Partial output before error" in content
|
||||
assert "Error:" in content
|
||||
assert "Unexpected termination" in content
|
||||
@@ -0,0 +1,62 @@
|
||||
"""Tests for Anthropic native tool base classes."""
|
||||
|
||||
import pytest
|
||||
from nanobot.agent.tools.anthropic.base import (
|
||||
BaseAnthropicTool,
|
||||
ToolResult,
|
||||
CLIResult,
|
||||
ToolError,
|
||||
)
|
||||
|
||||
|
||||
class DummyTool(BaseAnthropicTool):
|
||||
"""Test tool implementation."""
|
||||
api_type = "test_20250227"
|
||||
name = "test_tool"
|
||||
beta_flag = "test-beta"
|
||||
|
||||
async def __call__(self, **kwargs):
|
||||
return ToolResult(output="test output")
|
||||
|
||||
def to_params(self):
|
||||
return {"type": self.api_type, "name": self.name}
|
||||
|
||||
|
||||
def test_tool_result_dataclass():
|
||||
"""Test ToolResult can be created with all fields."""
|
||||
result = ToolResult(output="hello", error=None, base64_image=None, system="system message")
|
||||
assert result.output == "hello"
|
||||
assert result.error is None
|
||||
assert result.base64_image is None
|
||||
assert result.system == "system message"
|
||||
|
||||
|
||||
def test_cli_result_dataclass():
|
||||
"""Test CLIResult can be created with all fields."""
|
||||
result = CLIResult(exit_code=0, output="command output", error="")
|
||||
assert result.output == "command output"
|
||||
assert result.exit_code == 0
|
||||
assert result.error == ""
|
||||
|
||||
|
||||
def test_tool_error_exception():
|
||||
"""Test ToolError can be raised and caught."""
|
||||
with pytest.raises(ToolError):
|
||||
raise ToolError("Test error message")
|
||||
|
||||
|
||||
def test_base_anthropic_tool_to_params():
|
||||
"""Test tool returns correct params format."""
|
||||
tool = DummyTool()
|
||||
params = tool.to_params()
|
||||
assert params["type"] == "test_20250227"
|
||||
assert params["name"] == "test_tool"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_base_anthropic_tool_call():
|
||||
"""Test tool can be called and returns ToolResult."""
|
||||
tool = DummyTool()
|
||||
result = await tool()
|
||||
assert isinstance(result, ToolResult)
|
||||
assert result.output == "test output"
|
||||
@@ -0,0 +1,74 @@
|
||||
"""Tests for native tool support in AnthropicOAuthProvider."""
|
||||
|
||||
from nanobot.providers.anthropic_oauth import AnthropicOAuthProvider
|
||||
|
||||
|
||||
def test_convert_tools_passes_through_native_tools():
|
||||
"""Test that native tool format is passed through unchanged."""
|
||||
provider = AnthropicOAuthProvider(oauth_token="test", thinking_budget=0)
|
||||
|
||||
tools = [
|
||||
{
|
||||
"type": "bash_20250124",
|
||||
"name": "bash"
|
||||
}
|
||||
]
|
||||
|
||||
result = provider._convert_tools_to_anthropic(tools)
|
||||
assert len(result) == 1
|
||||
assert result[0]["type"] == "bash_20250124"
|
||||
assert result[0]["name"] == "bash"
|
||||
|
||||
|
||||
def test_convert_tools_handles_mixed_tool_types():
|
||||
"""Test conversion of both function and native tools."""
|
||||
provider = AnthropicOAuthProvider(oauth_token="test", thinking_budget=0)
|
||||
|
||||
tools = [
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "custom_tool",
|
||||
"description": "A custom tool",
|
||||
"parameters": {"type": "object", "properties": {"arg": {"type": "string"}}}
|
||||
}
|
||||
},
|
||||
{
|
||||
"type": "bash_20250124",
|
||||
"name": "bash"
|
||||
}
|
||||
]
|
||||
|
||||
result = provider._convert_tools_to_anthropic(tools)
|
||||
assert len(result) == 2
|
||||
|
||||
# Function tool gets converted
|
||||
assert result[0]["name"] == "custom_tool"
|
||||
assert result[0]["description"] == "A custom tool"
|
||||
assert "input_schema" in result[0]
|
||||
|
||||
# Native tool passed through
|
||||
assert result[1]["type"] == "bash_20250124"
|
||||
assert result[1]["name"] == "bash"
|
||||
|
||||
|
||||
def test_convert_tools_preserves_function_tool_conversion():
|
||||
"""Test that existing function tool conversion still works."""
|
||||
provider = AnthropicOAuthProvider(oauth_token="test", thinking_budget=0)
|
||||
|
||||
tools = [
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "test",
|
||||
"description": "desc",
|
||||
"parameters": {"type": "object"}
|
||||
}
|
||||
}
|
||||
]
|
||||
|
||||
result = provider._convert_tools_to_anthropic(tools)
|
||||
assert len(result) == 1
|
||||
assert result[0]["name"] == "test"
|
||||
assert result[0]["description"] == "desc"
|
||||
assert result[0]["input_schema"] == {"type": "object"}
|
||||
@@ -0,0 +1,56 @@
|
||||
"""Tests for BashTool20250124."""
|
||||
|
||||
import pytest
|
||||
from nanobot.agent.tools.anthropic.bash import BashTool20250124
|
||||
from nanobot.agent.tools.anthropic.base import ToolResult
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_bash_tool_simple_command():
|
||||
"""Test bash tool executes simple command."""
|
||||
tool = BashTool20250124()
|
||||
result = await tool(command="echo hello")
|
||||
|
||||
assert isinstance(result, ToolResult)
|
||||
assert "hello" in result.output
|
||||
assert result.error is None
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_bash_tool_persistent_session():
|
||||
"""Test bash tool maintains session across calls."""
|
||||
tool = BashTool20250124()
|
||||
|
||||
# Set variable
|
||||
result1 = await tool(command="export TEST_VAR=42")
|
||||
assert result1.error is None
|
||||
|
||||
# Read variable (should persist)
|
||||
result2 = await tool(command="echo $TEST_VAR")
|
||||
assert "42" in result2.output
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_bash_tool_restart():
|
||||
"""Test bash tool can restart session."""
|
||||
tool = BashTool20250124()
|
||||
|
||||
# Set variable
|
||||
await tool(command="export TEST_VAR=42")
|
||||
|
||||
# Restart
|
||||
result = await tool(restart=True)
|
||||
assert "restarted" in result.output.lower()
|
||||
|
||||
# Variable should be gone
|
||||
result2 = await tool(command="echo $TEST_VAR")
|
||||
assert "42" not in result2.output
|
||||
|
||||
|
||||
def test_bash_tool_to_params():
|
||||
"""Test bash tool returns correct params."""
|
||||
tool = BashTool20250124()
|
||||
params = tool.to_params()
|
||||
|
||||
assert params["type"] == "bash_20250124"
|
||||
assert params["name"] == "bash"
|
||||
@@ -0,0 +1,103 @@
|
||||
"""Tests for beta flag collection from native tools."""
|
||||
|
||||
from unittest.mock import AsyncMock, MagicMock, patch
|
||||
|
||||
import pytest
|
||||
|
||||
from nanobot.providers.anthropic_oauth import AnthropicOAuthProvider
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_beta_flags_collected_from_tools():
|
||||
"""Test that beta flags are extracted from tool objects."""
|
||||
provider = AnthropicOAuthProvider(oauth_token="test", thinking_budget=0)
|
||||
|
||||
# Mock tool objects with beta_flag attribute and to_params method
|
||||
class MockTool:
|
||||
def __init__(self, beta_flag):
|
||||
self.beta_flag = beta_flag
|
||||
|
||||
def to_params(self):
|
||||
return {"type": "bash_20250124", "name": "bash"}
|
||||
|
||||
tools_with_flags = [
|
||||
MockTool("computer-use-2025-11-24"),
|
||||
MockTool("computer-use-2025-11-24"), # Duplicate should be deduplicated
|
||||
]
|
||||
|
||||
# We need to test this via the actual API call flow
|
||||
# Mock httpx client
|
||||
mock_response = MagicMock()
|
||||
mock_response.status_code = 200
|
||||
mock_response.json.return_value = {
|
||||
"id": "msg_test",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [{"type": "text", "text": "test"}],
|
||||
"model": "claude-opus-4",
|
||||
"stop_reason": "end_turn",
|
||||
"usage": {"input_tokens": 10, "output_tokens": 10}
|
||||
}
|
||||
|
||||
with patch.object(provider, '_client') as mock_client:
|
||||
mock_client.post = AsyncMock(return_value=mock_response)
|
||||
|
||||
# Call with messages and tools
|
||||
await provider.chat(
|
||||
messages=[{"role": "user", "content": "test"}],
|
||||
model="claude-opus-4",
|
||||
max_tokens=100,
|
||||
tools=tools_with_flags
|
||||
)
|
||||
|
||||
# Check that beta flag was added to headers
|
||||
call_args = mock_client.post.call_args
|
||||
headers = call_args[1]["headers"]
|
||||
assert "anthropic-beta" in headers
|
||||
assert headers["anthropic-beta"] == "computer-use-2025-11-24"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_multiple_beta_flags_joined():
|
||||
"""Test that multiple unique beta flags are joined with commas."""
|
||||
provider = AnthropicOAuthProvider(oauth_token="test", thinking_budget=0)
|
||||
|
||||
class MockTool:
|
||||
def __init__(self, beta_flag):
|
||||
self.beta_flag = beta_flag
|
||||
|
||||
def to_params(self):
|
||||
return {"type": "bash_20250124", "name": "bash"}
|
||||
|
||||
tools_with_flags = [
|
||||
MockTool("flag-a"),
|
||||
MockTool("flag-b"),
|
||||
]
|
||||
|
||||
mock_response = MagicMock()
|
||||
mock_response.status_code = 200
|
||||
mock_response.json.return_value = {
|
||||
"id": "msg_test",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [{"type": "text", "text": "test"}],
|
||||
"model": "claude-opus-4",
|
||||
"stop_reason": "end_turn",
|
||||
"usage": {"input_tokens": 10, "output_tokens": 10}
|
||||
}
|
||||
|
||||
with patch.object(provider, '_client') as mock_client:
|
||||
mock_client.post = AsyncMock(return_value=mock_response)
|
||||
|
||||
await provider.chat(
|
||||
messages=[{"role": "user", "content": "test"}],
|
||||
model="claude-opus-4",
|
||||
max_tokens=100,
|
||||
tools=tools_with_flags
|
||||
)
|
||||
|
||||
call_args = mock_client.post.call_args
|
||||
headers = call_args[1]["headers"]
|
||||
assert "anthropic-beta" in headers
|
||||
# Should be sorted alphabetically and joined with comma
|
||||
assert headers["anthropic-beta"] == "flag-a,flag-b"
|
||||
@@ -0,0 +1,82 @@
|
||||
"""Tests for ComputerTool20251124."""
|
||||
|
||||
import pytest
|
||||
from unittest.mock import AsyncMock, patch, MagicMock
|
||||
from nanobot.agent.tools.anthropic.computer import ComputerTool20251124
|
||||
from nanobot.agent.tools.anthropic.base import ToolResult
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_computer_tool_screenshot():
|
||||
"""Test computer tool can take screenshot."""
|
||||
tool = ComputerTool20251124(vnc_host="localhost", vnc_port=5900)
|
||||
|
||||
# Mock VNC client
|
||||
with patch('nanobot.agent.tools.anthropic.computer.VNCDoToolClient') as mock_vnc:
|
||||
mock_client = AsyncMock()
|
||||
mock_client.captureScreen = AsyncMock(return_value=b"fake_png_data")
|
||||
|
||||
# Set up async context manager
|
||||
mock_context = MagicMock()
|
||||
mock_context.__aenter__ = AsyncMock(return_value=mock_client)
|
||||
mock_context.__aexit__ = AsyncMock(return_value=None)
|
||||
mock_vnc.create = MagicMock(return_value=mock_context)
|
||||
|
||||
result = await tool(action="screenshot")
|
||||
|
||||
assert isinstance(result, ToolResult)
|
||||
assert result.base64_image is not None
|
||||
assert len(result.base64_image) > 0
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_computer_tool_mouse_move():
|
||||
"""Test computer tool can move mouse."""
|
||||
tool = ComputerTool20251124(vnc_host="localhost", vnc_port=5900)
|
||||
|
||||
with patch('nanobot.agent.tools.anthropic.computer.VNCDoToolClient') as mock_vnc:
|
||||
mock_client = AsyncMock()
|
||||
mock_client.mouseMove = AsyncMock()
|
||||
|
||||
# Set up async context manager
|
||||
mock_context = MagicMock()
|
||||
mock_context.__aenter__ = AsyncMock(return_value=mock_client)
|
||||
mock_context.__aexit__ = AsyncMock(return_value=None)
|
||||
mock_vnc.create = MagicMock(return_value=mock_context)
|
||||
|
||||
result = await tool(action="mouse_move", coordinate=[100, 200])
|
||||
|
||||
assert isinstance(result, ToolResult)
|
||||
assert result.error is None
|
||||
mock_client.mouseMove.assert_called_once_with(100, 200)
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_computer_tool_key():
|
||||
"""Test computer tool can press keys."""
|
||||
tool = ComputerTool20251124(vnc_host="localhost", vnc_port=5900)
|
||||
|
||||
with patch('nanobot.agent.tools.anthropic.computer.VNCDoToolClient') as mock_vnc:
|
||||
mock_client = AsyncMock()
|
||||
mock_client.keyPress = AsyncMock()
|
||||
|
||||
# Set up async context manager
|
||||
mock_context = MagicMock()
|
||||
mock_context.__aenter__ = AsyncMock(return_value=mock_client)
|
||||
mock_context.__aexit__ = AsyncMock(return_value=None)
|
||||
mock_vnc.create = MagicMock(return_value=mock_context)
|
||||
|
||||
result = await tool(action="key", text="Return")
|
||||
|
||||
assert isinstance(result, ToolResult)
|
||||
assert result.error is None
|
||||
mock_client.keyPress.assert_called_once_with("Return")
|
||||
|
||||
|
||||
def test_computer_tool_to_params():
|
||||
"""Test computer tool returns correct params."""
|
||||
tool = ComputerTool20251124(vnc_host="localhost", vnc_port=5900)
|
||||
params = tool.to_params()
|
||||
|
||||
assert params["type"] == "computer_20251124"
|
||||
assert params["name"] == "computer"
|
||||
@@ -0,0 +1,116 @@
|
||||
"""Tests for EditTool20250728."""
|
||||
|
||||
import pytest
|
||||
from pathlib import Path
|
||||
from nanobot.agent.tools.anthropic.edit import EditTool20250728
|
||||
from nanobot.agent.tools.anthropic.base import CLIResult
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def edit_tool():
|
||||
"""Create an EditTool20250728 instance."""
|
||||
return EditTool20250728()
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def temp_file(tmp_path):
|
||||
"""Create a temporary file with some content."""
|
||||
file_path = tmp_path / "test.txt"
|
||||
file_path.write_text("line 1\nline 2\nline 3\n")
|
||||
return file_path
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_view_command(edit_tool, temp_file):
|
||||
"""Test viewing a file with line numbers."""
|
||||
result = await edit_tool(
|
||||
command="view",
|
||||
path=str(temp_file)
|
||||
)
|
||||
assert result.output is not None
|
||||
assert "1|line 1" in result.output
|
||||
assert "2|line 2" in result.output
|
||||
assert "3|line 3" in result.output
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_create_command(edit_tool, tmp_path):
|
||||
"""Test creating a new file."""
|
||||
new_file = tmp_path / "new.txt"
|
||||
result = await edit_tool(
|
||||
command="create",
|
||||
path=str(new_file),
|
||||
file_text="Hello\nWorld\n"
|
||||
)
|
||||
assert result.exit_code == 0
|
||||
assert new_file.exists()
|
||||
assert new_file.read_text() == "Hello\nWorld\n"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_str_replace_command(edit_tool, temp_file):
|
||||
"""Test replacing a unique string."""
|
||||
result = await edit_tool(
|
||||
command="str_replace",
|
||||
path=str(temp_file),
|
||||
old_str="line 2",
|
||||
new_str="LINE TWO"
|
||||
)
|
||||
assert result.exit_code == 0
|
||||
content = temp_file.read_text()
|
||||
assert "LINE TWO" in content
|
||||
assert "line 2" not in content
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_str_replace_non_unique(edit_tool, temp_file):
|
||||
"""Test that str_replace fails on non-unique match."""
|
||||
# Write content with duplicate "line"
|
||||
temp_file.write_text("line 1\nline 2\nline 3\n")
|
||||
result = await edit_tool(
|
||||
command="str_replace",
|
||||
path=str(temp_file),
|
||||
old_str="line", # This appears 3 times
|
||||
new_str="LINE"
|
||||
)
|
||||
assert result.exit_code == 1
|
||||
assert "must match exactly once" in result.error.lower()
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_insert_command(edit_tool, temp_file):
|
||||
"""Test inserting text at a specific line."""
|
||||
result = await edit_tool(
|
||||
command="insert",
|
||||
path=str(temp_file),
|
||||
insert_line=1,
|
||||
new_str="inserted line\n"
|
||||
)
|
||||
assert result.exit_code == 0
|
||||
content = temp_file.read_text()
|
||||
lines = content.splitlines()
|
||||
assert lines[1] == "inserted line"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_edit_tool_requires_absolute_path():
|
||||
"""Test edit tool rejects relative paths."""
|
||||
tool = EditTool20250728()
|
||||
|
||||
result = await tool(
|
||||
command="view",
|
||||
path="relative/path.txt"
|
||||
)
|
||||
|
||||
assert isinstance(result, CLIResult)
|
||||
assert result.exit_code == 1
|
||||
assert "absolute" in result.error.lower()
|
||||
|
||||
|
||||
def test_edit_tool_to_params():
|
||||
"""Test edit tool returns correct params."""
|
||||
tool = EditTool20250728()
|
||||
params = tool.to_params()
|
||||
|
||||
assert params["type"] == "text_editor_20250728"
|
||||
assert params["name"] == "str_replace_editor"
|
||||
@@ -0,0 +1,146 @@
|
||||
# tests/test_heartbeat_idle_detection.py
|
||||
"""Tests for heartbeat idle detection with sender_id filtering."""
|
||||
import pytest
|
||||
from datetime import datetime, timedelta
|
||||
from pathlib import Path
|
||||
from nanobot.heartbeat.service import HeartbeatService
|
||||
from nanobot.session.manager import Session, SessionManager
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_idle_detection_ignores_system_messages():
|
||||
"""Test that heartbeat only counts real user messages for idle detection."""
|
||||
# Create session manager and session
|
||||
session_manager = SessionManager(workspace=Path("/tmp/test-heartbeat"))
|
||||
session = session_manager.get_or_create("telegram:239824268")
|
||||
|
||||
# Add a real user message 45 minutes ago
|
||||
real_user_time = datetime.now() - timedelta(minutes=45)
|
||||
session.add_message(
|
||||
"user",
|
||||
"This is a real user message",
|
||||
sender_id="239824268|testuser",
|
||||
timestamp=real_user_time.isoformat()
|
||||
)
|
||||
|
||||
# Add a heartbeat system message 10 minutes ago (should be ignored)
|
||||
heartbeat_time = datetime.now() - timedelta(minutes=10)
|
||||
session.add_message(
|
||||
"user",
|
||||
"Read HEARTBEAT.md...",
|
||||
sender_id="user",
|
||||
timestamp=heartbeat_time.isoformat()
|
||||
)
|
||||
|
||||
# Add an assistant response
|
||||
session.add_message("assistant", "Response to heartbeat")
|
||||
|
||||
# Create heartbeat service with 30 minute idle threshold
|
||||
heartbeat = HeartbeatService(
|
||||
workspace=Path("/tmp/test-heartbeat"),
|
||||
session_manager=session_manager,
|
||||
target_session_key="telegram:239824268",
|
||||
idle_threshold_s=30 * 60, # 30 minutes
|
||||
interval_s=30 * 60,
|
||||
enabled=False # Don't actually start the loop
|
||||
)
|
||||
|
||||
# Manually check idle logic (replicate _tick logic)
|
||||
session = session_manager.get_or_create("telegram:239824268")
|
||||
|
||||
# Find last user message timestamp (should find the 45-minute-old message, not the 10-minute-old one)
|
||||
last_user_timestamp = None
|
||||
for msg in reversed(session.messages):
|
||||
if msg.get("role") == "user":
|
||||
sender_id = msg.get("sender_id")
|
||||
if sender_id == "user":
|
||||
continue
|
||||
last_user_timestamp = msg.get("timestamp")
|
||||
break
|
||||
|
||||
assert last_user_timestamp is not None
|
||||
last_dt = datetime.fromisoformat(last_user_timestamp)
|
||||
elapsed = (datetime.now() - last_dt).total_seconds()
|
||||
|
||||
# Should detect user is idle (45 minutes > 30 minute threshold)
|
||||
assert elapsed >= 30 * 60, f"Expected idle (45min), but elapsed={elapsed/60:.1f}min"
|
||||
# Should NOT be 10 minutes (heartbeat message was ignored)
|
||||
assert elapsed >= 40 * 60, f"Heartbeat message was not ignored, elapsed={elapsed/60:.1f}min"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_idle_detection_counts_real_user_messages():
|
||||
"""Test that heartbeat correctly identifies when user is active."""
|
||||
# Create session manager and session
|
||||
session_manager = SessionManager(workspace=Path("/tmp/test-heartbeat"))
|
||||
session = session_manager.get_or_create("telegram:239824268")
|
||||
|
||||
# Add a real user message 10 minutes ago (recent activity)
|
||||
real_user_time = datetime.now() - timedelta(minutes=10)
|
||||
session.add_message(
|
||||
"user",
|
||||
"This is a recent user message",
|
||||
sender_id="239824268|testuser",
|
||||
timestamp=real_user_time.isoformat()
|
||||
)
|
||||
|
||||
# Create heartbeat service with 30 minute idle threshold
|
||||
heartbeat = HeartbeatService(
|
||||
workspace=Path("/tmp/test-heartbeat"),
|
||||
session_manager=session_manager,
|
||||
target_session_key="telegram:239824268",
|
||||
idle_threshold_s=30 * 60, # 30 minutes
|
||||
interval_s=30 * 60,
|
||||
enabled=False
|
||||
)
|
||||
|
||||
# Find last user message timestamp
|
||||
session = session_manager.get_or_create("telegram:239824268")
|
||||
last_user_timestamp = None
|
||||
for msg in reversed(session.messages):
|
||||
if msg.get("role") == "user":
|
||||
sender_id = msg.get("sender_id")
|
||||
if sender_id == "user":
|
||||
continue
|
||||
last_user_timestamp = msg.get("timestamp")
|
||||
break
|
||||
|
||||
assert last_user_timestamp is not None
|
||||
last_dt = datetime.fromisoformat(last_user_timestamp)
|
||||
elapsed = (datetime.now() - last_dt).total_seconds()
|
||||
|
||||
# Should detect user is active (10 minutes < 30 minute threshold)
|
||||
assert elapsed < 30 * 60, f"Expected active (10min), but elapsed={elapsed/60:.1f}min"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_backwards_compat_messages_without_sender_id():
|
||||
"""Test that old messages without sender_id are treated as real user messages."""
|
||||
# Create session manager and session
|
||||
session_manager = SessionManager(workspace=Path("/tmp/test-heartbeat"))
|
||||
session = session_manager.get_or_create("telegram:239824268")
|
||||
|
||||
# Add an old message without sender_id (backwards compat)
|
||||
old_time = datetime.now() - timedelta(minutes=20)
|
||||
session.add_message(
|
||||
"user",
|
||||
"Old message without sender_id",
|
||||
timestamp=old_time.isoformat()
|
||||
)
|
||||
|
||||
# Find last user message timestamp (should find the old message)
|
||||
last_user_timestamp = None
|
||||
for msg in reversed(session.messages):
|
||||
if msg.get("role") == "user":
|
||||
sender_id = msg.get("sender_id")
|
||||
if sender_id == "user":
|
||||
continue
|
||||
last_user_timestamp = msg.get("timestamp")
|
||||
break
|
||||
|
||||
assert last_user_timestamp is not None
|
||||
last_dt = datetime.fromisoformat(last_user_timestamp)
|
||||
elapsed = (datetime.now() - last_dt).total_seconds()
|
||||
|
||||
# Should accept old message (backwards compat)
|
||||
assert elapsed < 25 * 60, f"Old message not counted, elapsed={elapsed/60:.1f}min"
|
||||
@@ -15,7 +15,7 @@ from unittest.mock import AsyncMock, MagicMock
|
||||
async def test_idle_heartbeat_end_to_end(tmp_path):
|
||||
"""
|
||||
Integration test: heartbeat triggers when idle, runs in main session,
|
||||
output is suppressed, session contains [HIDDEN] content.
|
||||
output is suppressed, session contains [HIDDEN:signature] content.
|
||||
"""
|
||||
workspace = tmp_path / "test-integration"
|
||||
workspace.mkdir()
|
||||
@@ -91,14 +91,20 @@ async def test_idle_heartbeat_end_to_end(tmp_path):
|
||||
# 1. Session has new messages
|
||||
assert len(session.messages) > 1
|
||||
|
||||
# 2. Find the heartbeat response (assistant message)
|
||||
# 2. Find the heartbeat response (assistant message with signed visibility marker)
|
||||
heartbeat_messages = [
|
||||
m for m in session.messages
|
||||
if m.get("role") == "assistant" and "[HIDDEN]" in m.get("content", "")
|
||||
if m.get("role") == "assistant" and "[HIDDEN:" in m.get("content", "")
|
||||
]
|
||||
assert len(heartbeat_messages) == 1, "Expected exactly 1 [HIDDEN] heartbeat message"
|
||||
assert len(heartbeat_messages) == 1, "Expected exactly 1 [HIDDEN:signature] heartbeat message"
|
||||
|
||||
# 3. Verify content is prefixed with [HIDDEN]
|
||||
# 3. Verify content is prefixed with [HIDDEN:signature]
|
||||
heartbeat_msg = heartbeat_messages[0]
|
||||
assert heartbeat_msg["content"].startswith("[HIDDEN]")
|
||||
assert heartbeat_msg["content"].startswith("[HIDDEN:")
|
||||
# Verify signature format (8-char hex)
|
||||
content = heartbeat_msg["content"]
|
||||
prefix_end = content.index("]")
|
||||
signature = content[8:prefix_end] # Skip "[HIDDEN:" to get signature
|
||||
assert len(signature) == 8, f"Expected 8-char signature, got {len(signature)}"
|
||||
assert all(c in "0123456789abcdef" for c in signature), "Signature should be hex"
|
||||
assert "Heartbeat executed successfully" in heartbeat_msg["content"]
|
||||
|
||||
@@ -0,0 +1,25 @@
|
||||
"""Tests for screenshot media tracking."""
|
||||
|
||||
import pytest
|
||||
import base64
|
||||
from pathlib import Path
|
||||
from unittest.mock import AsyncMock, MagicMock
|
||||
from nanobot.agent.loop import AgentLoop
|
||||
from nanobot.agent.tools.anthropic.base import ToolResult
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_media_tracking_saves_screenshots():
|
||||
"""Test that screenshots are saved to disk and tracked."""
|
||||
# This is more of an integration test
|
||||
# Test the media saving logic separately
|
||||
|
||||
# Create fake screenshot data
|
||||
fake_png = b"\x89PNG\r\n\x1a\n" # PNG header
|
||||
base64_image = base64.b64encode(fake_png).decode()
|
||||
|
||||
result = ToolResult(base64_image=base64_image)
|
||||
|
||||
# Verify we can decode it
|
||||
decoded = base64.b64decode(result.base64_image)
|
||||
assert decoded == fake_png
|
||||
@@ -0,0 +1,57 @@
|
||||
"""Test registration of native Anthropic tools in the agent loop."""
|
||||
|
||||
import pytest
|
||||
from pathlib import Path
|
||||
from unittest.mock import AsyncMock, MagicMock
|
||||
|
||||
from nanobot.agent.loop import AgentLoop
|
||||
from nanobot.agent.tools.anthropic import (
|
||||
BashTool20250124,
|
||||
EditTool20250728,
|
||||
ComputerTool20251124,
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def mock_provider():
|
||||
"""Create a mock provider."""
|
||||
provider = MagicMock()
|
||||
provider.chat = AsyncMock(return_value="test response")
|
||||
provider.get_default_model = MagicMock(return_value="test-model")
|
||||
return provider
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def mock_bus():
|
||||
"""Create a mock message bus."""
|
||||
bus = MagicMock()
|
||||
bus.publish_outbound = AsyncMock()
|
||||
return bus
|
||||
|
||||
|
||||
def test_native_tools_registered(mock_provider, mock_bus, tmp_path):
|
||||
"""Test that native Anthropic tools are registered in the agent loop."""
|
||||
# Create agent loop
|
||||
loop = AgentLoop(
|
||||
provider=mock_provider,
|
||||
bus=mock_bus,
|
||||
workspace=tmp_path,
|
||||
)
|
||||
|
||||
# Get all registered tool names
|
||||
tool_names = [tool.name for tool in loop.tools._tools.values()]
|
||||
|
||||
# Verify native tools are registered (using their internal names)
|
||||
assert "bash" in tool_names, "bash tool should be registered"
|
||||
assert "str_replace_editor" in tool_names, "str_replace_editor tool should be registered"
|
||||
assert "computer" in tool_names, "computer tool should be registered"
|
||||
|
||||
# Verify we can get the tool instances
|
||||
bash_tool = loop.tools.get("bash")
|
||||
assert isinstance(bash_tool, BashTool20250124)
|
||||
|
||||
editor_tool = loop.tools.get("str_replace_editor")
|
||||
assert isinstance(editor_tool, EditTool20250728)
|
||||
|
||||
computer_tool = loop.tools.get("computer")
|
||||
assert isinstance(computer_tool, ComputerTool20251124)
|
||||
@@ -0,0 +1,94 @@
|
||||
"""Tests for registry duck typing support."""
|
||||
|
||||
import pytest
|
||||
from nanobot.agent.tools.registry import ToolRegistry
|
||||
from nanobot.agent.tools.anthropic.base import BaseAnthropicTool, ToolResult
|
||||
|
||||
|
||||
class MockNativeTool(BaseAnthropicTool):
|
||||
"""Mock native tool for testing."""
|
||||
api_type = "test_20250227"
|
||||
name = "native_test"
|
||||
beta_flag = "test-beta"
|
||||
|
||||
async def __call__(self, **kwargs):
|
||||
return ToolResult(output="native result")
|
||||
|
||||
def to_params(self):
|
||||
return {"type": self.api_type, "name": self.name}
|
||||
|
||||
|
||||
class MockFunctionTool:
|
||||
"""Mock function tool for testing."""
|
||||
def __init__(self):
|
||||
self.name = "function_test"
|
||||
|
||||
def to_schema(self):
|
||||
return {
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": self.name,
|
||||
"description": "Test function tool",
|
||||
"parameters": {"type": "object", "properties": {}}
|
||||
}
|
||||
}
|
||||
|
||||
async def execute(self, **kwargs):
|
||||
return "function result"
|
||||
|
||||
|
||||
def test_registry_supports_native_tools():
|
||||
"""Test registry can register and get definitions from native tools."""
|
||||
registry = ToolRegistry()
|
||||
native_tool = MockNativeTool()
|
||||
registry.register(native_tool)
|
||||
|
||||
definitions = registry.get_definitions()
|
||||
assert len(definitions) == 1
|
||||
assert definitions[0]["type"] == "test_20250227"
|
||||
assert definitions[0]["name"] == "native_test"
|
||||
|
||||
|
||||
def test_registry_supports_function_tools():
|
||||
"""Test registry still supports function tools."""
|
||||
registry = ToolRegistry()
|
||||
function_tool = MockFunctionTool()
|
||||
registry.register(function_tool)
|
||||
|
||||
definitions = registry.get_definitions()
|
||||
assert len(definitions) == 1
|
||||
assert definitions[0]["type"] == "function"
|
||||
assert definitions[0]["function"]["name"] == "function_test"
|
||||
|
||||
|
||||
def test_registry_supports_mixed_tools():
|
||||
"""Test registry can handle both native and function tools."""
|
||||
registry = ToolRegistry()
|
||||
native_tool = MockNativeTool()
|
||||
function_tool = MockFunctionTool()
|
||||
|
||||
registry.register(native_tool)
|
||||
registry.register(function_tool)
|
||||
|
||||
definitions = registry.get_definitions()
|
||||
assert len(definitions) == 2
|
||||
|
||||
# Find each tool type in definitions
|
||||
native_def = next(d for d in definitions if d.get("type") == "test_20250227")
|
||||
function_def = next(d for d in definitions if d.get("type") == "function")
|
||||
|
||||
assert native_def["name"] == "native_test"
|
||||
assert function_def["function"]["name"] == "function_test"
|
||||
|
||||
|
||||
def test_registry_rejects_tools_without_schema_method():
|
||||
"""Test registry raises error for tools with no schema method."""
|
||||
registry = ToolRegistry()
|
||||
|
||||
class BadTool:
|
||||
name = "bad"
|
||||
|
||||
registry.register(BadTool())
|
||||
|
||||
with pytest.raises(ValueError, match="has no schema method"):
|
||||
registry.get_definitions()
|
||||
@@ -0,0 +1,57 @@
|
||||
"""Tests for registry execution of native tools."""
|
||||
|
||||
import pytest
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
from nanobot.agent.tools.registry import ToolRegistry
|
||||
from nanobot.agent.tools.anthropic import BashTool20250124, EditTool20250728
|
||||
from nanobot.agent.tools.anthropic.base import ToolResult, CLIResult
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_registry_executes_bash_tool():
|
||||
"""Test registry can execute BashTool20250124 and returns ToolResult."""
|
||||
registry = ToolRegistry()
|
||||
registry.register(BashTool20250124())
|
||||
|
||||
result = await registry.execute("bash", {"command": "echo 'test'"})
|
||||
|
||||
assert isinstance(result, ToolResult)
|
||||
assert result.output is not None
|
||||
assert "test" in result.output
|
||||
assert result.error is None
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_registry_executes_edit_tool():
|
||||
"""Test registry can execute EditTool20250728 and returns CLIResult."""
|
||||
registry = ToolRegistry()
|
||||
registry.register(EditTool20250728())
|
||||
|
||||
with tempfile.TemporaryDirectory() as tmpdir:
|
||||
test_file = str(Path(tmpdir) / "test.txt")
|
||||
|
||||
result = await registry.execute("str_replace_editor", {
|
||||
"command": "create",
|
||||
"path": test_file,
|
||||
"file_text": "Hello, world!"
|
||||
})
|
||||
|
||||
assert isinstance(result, CLIResult)
|
||||
assert "created" in result.output.lower() or "success" in result.output.lower()
|
||||
assert Path(test_file).exists()
|
||||
assert Path(test_file).read_text() == "Hello, world!"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_registry_mixed_tools():
|
||||
"""Test registry can execute both native and function tools in same registry."""
|
||||
registry = ToolRegistry()
|
||||
|
||||
# Register native tool
|
||||
registry.register(BashTool20250124())
|
||||
|
||||
# Execute native tool
|
||||
result = await registry.execute("bash", {"command": "echo 'native'"})
|
||||
assert isinstance(result, ToolResult)
|
||||
assert "native" in result.output
|
||||
@@ -0,0 +1,101 @@
|
||||
"""Integration tests for Telegram media sending."""
|
||||
|
||||
from unittest.mock import AsyncMock, MagicMock, patch
|
||||
|
||||
import pytest
|
||||
|
||||
from nanobot.bus.events import OutboundMessage
|
||||
from nanobot.bus.queue import MessageBus
|
||||
from nanobot.channels.telegram import TelegramChannel
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def mock_telegram_app():
|
||||
"""Mock python-telegram-bot Application."""
|
||||
app = MagicMock()
|
||||
app.bot = MagicMock()
|
||||
app.bot.send_photo = AsyncMock()
|
||||
app.bot.send_video = AsyncMock()
|
||||
app.bot.send_audio = AsyncMock()
|
||||
app.bot.send_document = AsyncMock()
|
||||
app.bot.send_media_group = AsyncMock()
|
||||
app.bot.send_message = AsyncMock()
|
||||
return app
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_send_single_image(mock_telegram_app, tmp_path):
|
||||
"""Test sending single image."""
|
||||
# Create test image
|
||||
from PIL import Image
|
||||
img_path = tmp_path / "test.jpg"
|
||||
img = Image.new("RGB", (100, 100), color="red")
|
||||
img.save(img_path, format="JPEG")
|
||||
|
||||
# Setup channel
|
||||
bus = MessageBus()
|
||||
config = MagicMock()
|
||||
config.token = "fake_token"
|
||||
config.proxy = None
|
||||
|
||||
channel = TelegramChannel(config, bus)
|
||||
channel._app = mock_telegram_app
|
||||
channel._running = True
|
||||
|
||||
# Send message with media
|
||||
msg = OutboundMessage(
|
||||
channel="telegram",
|
||||
chat_id="12345",
|
||||
content="Test image",
|
||||
media=[str(img_path)]
|
||||
)
|
||||
|
||||
with patch("nanobot.channels.telegram._markdown_to_telegram_html", return_value="Test image"):
|
||||
await channel.send(msg)
|
||||
|
||||
# Verify send_photo was called
|
||||
mock_telegram_app.bot.send_photo.assert_called_once()
|
||||
call_args = mock_telegram_app.bot.send_photo.call_args
|
||||
assert call_args.kwargs["chat_id"] == 12345
|
||||
assert call_args.kwargs["caption"] == "Test image"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_send_album(mock_telegram_app, tmp_path):
|
||||
"""Test sending multiple images as album."""
|
||||
from PIL import Image
|
||||
|
||||
# Create test images
|
||||
img_paths = []
|
||||
for i in range(3):
|
||||
img_path = tmp_path / f"test{i}.jpg"
|
||||
img = Image.new("RGB", (100, 100), color="red")
|
||||
img.save(img_path, format="JPEG")
|
||||
img_paths.append(str(img_path))
|
||||
|
||||
# Setup channel
|
||||
bus = MessageBus()
|
||||
config = MagicMock()
|
||||
config.token = "fake_token"
|
||||
config.proxy = None
|
||||
|
||||
channel = TelegramChannel(config, bus)
|
||||
channel._app = mock_telegram_app
|
||||
channel._running = True
|
||||
|
||||
# Send message with multiple images
|
||||
msg = OutboundMessage(
|
||||
channel="telegram",
|
||||
chat_id="12345",
|
||||
content="Album test",
|
||||
media=img_paths
|
||||
)
|
||||
|
||||
with patch("nanobot.channels.telegram._markdown_to_telegram_html", return_value="Album test"):
|
||||
await channel.send(msg)
|
||||
|
||||
# Verify send_media_group was called
|
||||
mock_telegram_app.bot.send_media_group.assert_called_once()
|
||||
call_args = mock_telegram_app.bot.send_media_group.call_args
|
||||
assert call_args.kwargs["chat_id"] == 12345
|
||||
assert len(call_args.kwargs["media"]) == 3
|
||||
@@ -0,0 +1,263 @@
|
||||
"""Tests for Telegram media handling."""
|
||||
|
||||
import io
|
||||
|
||||
import pytest
|
||||
from PIL import Image
|
||||
|
||||
|
||||
def test_detect_mime_from_jpeg():
|
||||
"""Test MIME detection for JPEG images."""
|
||||
from nanobot.channels.telegram_media import detect_mime
|
||||
|
||||
# Create minimal JPEG bytes (FF D8 FF = JPEG magic bytes)
|
||||
jpeg_bytes = b'\xff\xd8\xff\xe0\x00\x10JFIF'
|
||||
|
||||
mime = detect_mime("test.jpg", jpeg_bytes)
|
||||
assert mime == "image/jpeg"
|
||||
|
||||
|
||||
def test_detect_mime_from_png():
|
||||
"""Test MIME detection for PNG images."""
|
||||
from nanobot.channels.telegram_media import detect_mime
|
||||
|
||||
# PNG magic bytes
|
||||
png_bytes = b'\x89PNG\r\n\x1a\n'
|
||||
|
||||
mime = detect_mime("test.png", png_bytes)
|
||||
assert mime == "image/png"
|
||||
|
||||
|
||||
def test_detect_mime_from_extension_fallback():
|
||||
"""Test MIME detection falls back to extension when no content provided."""
|
||||
from nanobot.channels.telegram_media import detect_mime
|
||||
|
||||
mime = detect_mime("video.mp4", None)
|
||||
assert mime == "video/mp4"
|
||||
|
||||
|
||||
def test_detect_mime_unknown():
|
||||
"""Test MIME detection returns generic type for unknown files."""
|
||||
from nanobot.channels.telegram_media import detect_mime
|
||||
|
||||
mime = detect_mime("unknown.xyz", None)
|
||||
assert mime == "application/octet-stream"
|
||||
|
||||
|
||||
def test_detect_mime_magic_fallback_on_octet_stream():
|
||||
"""Test that extension is preferred when magic returns generic type."""
|
||||
from nanobot.channels.telegram_media import detect_mime
|
||||
|
||||
# Generic binary content that magic might identify as octet-stream
|
||||
generic_bytes = b'\x00\x01\x02\x03'
|
||||
|
||||
# But extension clearly indicates it's an image
|
||||
mime = detect_mime("image.png", generic_bytes)
|
||||
|
||||
# Should use extension (png) not magic's generic result
|
||||
# Note: This tests the logic at line 36 - avoiding generic types
|
||||
assert mime in ("image/png", "application/octet-stream")
|
||||
|
||||
|
||||
def test_detect_mime_malformed_content():
|
||||
"""Test fallback when magic detection fails with malformed content."""
|
||||
from nanobot.channels.telegram_media import detect_mime
|
||||
|
||||
# Malformed content that might cause magic to raise an exception
|
||||
malformed = b'\xff' * 10
|
||||
|
||||
# Should fallback to extension detection, not crash
|
||||
mime = detect_mime("test.mp4", malformed)
|
||||
assert mime == "video/mp4"
|
||||
|
||||
|
||||
def test_classify_media_image():
|
||||
"""Test classification of image MIME types."""
|
||||
from nanobot.channels.telegram_media import MediaKind, classify_media
|
||||
|
||||
assert classify_media("image/jpeg") == MediaKind.IMAGE
|
||||
assert classify_media("image/png") == MediaKind.IMAGE
|
||||
assert classify_media("image/webp") == MediaKind.IMAGE
|
||||
|
||||
|
||||
def test_classify_media_video():
|
||||
"""Test classification of video MIME types."""
|
||||
from nanobot.channels.telegram_media import MediaKind, classify_media
|
||||
|
||||
assert classify_media("video/mp4") == MediaKind.VIDEO
|
||||
assert classify_media("video/quicktime") == MediaKind.VIDEO
|
||||
|
||||
|
||||
def test_classify_media_audio():
|
||||
"""Test classification of audio MIME types."""
|
||||
from nanobot.channels.telegram_media import MediaKind, classify_media
|
||||
|
||||
assert classify_media("audio/mpeg") == MediaKind.AUDIO
|
||||
assert classify_media("audio/ogg") == MediaKind.AUDIO
|
||||
|
||||
|
||||
def test_classify_media_document():
|
||||
"""Test classification of document MIME types."""
|
||||
from nanobot.channels.telegram_media import MediaKind, classify_media
|
||||
|
||||
assert classify_media("application/pdf") == MediaKind.DOCUMENT
|
||||
assert classify_media("text/plain") == MediaKind.DOCUMENT
|
||||
assert classify_media("application/octet-stream") == MediaKind.DOCUMENT
|
||||
|
||||
|
||||
def test_is_heic_format():
|
||||
"""Test HEIC format detection."""
|
||||
from nanobot.channels.telegram_media import is_heic_format
|
||||
|
||||
assert is_heic_format("photo.heic") is True
|
||||
assert is_heic_format("photo.HEIC") is True
|
||||
assert is_heic_format("photo.heif") is True
|
||||
assert is_heic_format("photo.jpg") is False
|
||||
|
||||
|
||||
def test_optimize_image_jpeg_quality(tmp_path):
|
||||
"""Test JPEG optimization reduces size with quality ladder."""
|
||||
from nanobot.channels.telegram_media import optimize_image
|
||||
|
||||
# Create a large test image (3000x3000 RGB)
|
||||
img = Image.new("RGB", (3000, 3000), color="red")
|
||||
buf = io.BytesIO()
|
||||
img.save(buf, format="JPEG", quality=95)
|
||||
original_bytes = buf.getvalue()
|
||||
original_size = len(original_bytes)
|
||||
|
||||
# Write to temp file
|
||||
temp_file = tmp_path / "test.jpg"
|
||||
temp_file.write_bytes(original_bytes)
|
||||
|
||||
# Optimize to 1MB max
|
||||
optimized = optimize_image(str(temp_file), max_bytes=1_000_000)
|
||||
|
||||
# Should be smaller than original
|
||||
assert len(optimized) < original_size
|
||||
# Should be under limit
|
||||
assert len(optimized) <= 1_000_000
|
||||
# Should still be valid JPEG
|
||||
assert optimized.startswith(b'\xff\xd8\xff')
|
||||
|
||||
|
||||
def test_optimize_image_png_preserve_alpha(tmp_path):
|
||||
"""Test PNG with alpha channel is preserved."""
|
||||
from nanobot.channels.telegram_media import optimize_image
|
||||
|
||||
# Create PNG with alpha channel
|
||||
img = Image.new("RGBA", (1000, 1000), color=(255, 0, 0, 128))
|
||||
buf = io.BytesIO()
|
||||
img.save(buf, format="PNG")
|
||||
original_bytes = buf.getvalue()
|
||||
|
||||
# Write to temp file
|
||||
temp_file = tmp_path / "test.png"
|
||||
temp_file.write_bytes(original_bytes)
|
||||
|
||||
optimized = optimize_image(str(temp_file), max_bytes=5_000_000)
|
||||
|
||||
# Should still be PNG (PNG magic bytes)
|
||||
assert optimized.startswith(b'\x89PNG')
|
||||
|
||||
# Load and verify alpha channel preserved
|
||||
img_opt = Image.open(io.BytesIO(optimized))
|
||||
assert img_opt.mode == "RGBA"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_fetch_media_success():
|
||||
"""Test fetching media from remote URL."""
|
||||
from unittest.mock import AsyncMock, MagicMock, patch
|
||||
|
||||
from nanobot.channels.telegram_media import fetch_media
|
||||
|
||||
# Mock httpx response
|
||||
mock_content = b"fake image data"
|
||||
mock_response = MagicMock()
|
||||
mock_response.content = mock_content
|
||||
mock_response.headers = {"content-type": "image/jpeg"}
|
||||
mock_response.raise_for_status = MagicMock()
|
||||
|
||||
with patch("httpx.AsyncClient") as mock_client:
|
||||
mock_client.return_value.__aenter__.return_value.get = AsyncMock(return_value=mock_response)
|
||||
|
||||
content, mime = await fetch_media("https://example.com/image.jpg", max_bytes=10_000_000)
|
||||
|
||||
assert content == mock_content
|
||||
assert mime == "image/jpeg"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_fetch_media_timeout():
|
||||
"""Test fetch media handles timeout."""
|
||||
from unittest.mock import AsyncMock, patch
|
||||
|
||||
import httpx
|
||||
|
||||
from nanobot.channels.telegram_media import fetch_media
|
||||
|
||||
with patch("httpx.AsyncClient") as mock_client:
|
||||
mock_client.return_value.__aenter__.return_value.get = AsyncMock(side_effect=httpx.TimeoutException("timeout"))
|
||||
|
||||
with pytest.raises(ValueError, match="timeout"):
|
||||
await fetch_media("https://example.com/image.jpg", max_bytes=10_000_000)
|
||||
|
||||
|
||||
def test_group_media_all_images():
|
||||
"""Test grouping all images into album."""
|
||||
from nanobot.channels.telegram_media import MediaKind, group_media_for_album
|
||||
|
||||
media_items = [
|
||||
("image1.jpg", MediaKind.IMAGE),
|
||||
("image2.png", MediaKind.IMAGE),
|
||||
("image3.jpeg", MediaKind.IMAGE),
|
||||
]
|
||||
|
||||
result = group_media_for_album(media_items)
|
||||
|
||||
assert result["album"] == ["image1.jpg", "image2.png", "image3.jpeg"]
|
||||
assert result["separate"] == []
|
||||
|
||||
|
||||
def test_group_media_all_videos():
|
||||
"""Test grouping all videos into album."""
|
||||
from nanobot.channels.telegram_media import MediaKind, group_media_for_album
|
||||
|
||||
media_items = [
|
||||
("video1.mp4", MediaKind.VIDEO),
|
||||
("video2.mov", MediaKind.VIDEO),
|
||||
]
|
||||
|
||||
result = group_media_for_album(media_items)
|
||||
|
||||
assert result["album"] == ["video1.mp4", "video2.mov"]
|
||||
assert result["separate"] == []
|
||||
|
||||
|
||||
def test_group_media_mixed_types():
|
||||
"""Test mixed media types sent separately."""
|
||||
from nanobot.channels.telegram_media import MediaKind, group_media_for_album
|
||||
|
||||
media_items = [
|
||||
("image.jpg", MediaKind.IMAGE),
|
||||
("video.mp4", MediaKind.VIDEO),
|
||||
("audio.mp3", MediaKind.AUDIO),
|
||||
]
|
||||
|
||||
result = group_media_for_album(media_items)
|
||||
|
||||
assert result["album"] == []
|
||||
assert result["separate"] == ["image.jpg", "video.mp4", "audio.mp3"]
|
||||
|
||||
|
||||
def test_group_media_single_item():
|
||||
"""Test single media item sent separately (not as album)."""
|
||||
from nanobot.channels.telegram_media import MediaKind, group_media_for_album
|
||||
|
||||
media_items = [("image.jpg", MediaKind.IMAGE)]
|
||||
|
||||
result = group_media_for_album(media_items)
|
||||
|
||||
assert result["album"] == []
|
||||
assert result["separate"] == ["image.jpg"]
|
||||
@@ -73,3 +73,433 @@ def test_strip_all_hidden_markers_removes_markers():
|
||||
forged = "[HIDDEN:deadbeef] Message"
|
||||
stripped = strip_all_hidden_markers(forged)
|
||||
assert stripped == "Message"
|
||||
|
||||
def test_system_prompt_includes_visibility_docs(tmp_path):
|
||||
"""Test that system prompt documents visibility markers."""
|
||||
from nanobot.agent.context import ContextBuilder
|
||||
|
||||
builder = ContextBuilder(workspace=tmp_path)
|
||||
prompt = builder.build_system_prompt()
|
||||
|
||||
# Should document visibility markers
|
||||
assert "[HIDDEN:" in prompt
|
||||
assert "cryptographically signed" in prompt.lower()
|
||||
assert "do not generate" in prompt.lower() or "don't generate" in prompt.lower()
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_suppress_mode_adds_signed_marker(tmp_path):
|
||||
"""Test that suppress mode adds cryptographically signed markers."""
|
||||
from unittest.mock import AsyncMock, Mock
|
||||
from nanobot.agent.loop import AgentLoop
|
||||
from nanobot.bus.queue import MessageBus
|
||||
from nanobot.session.manager import SessionManager
|
||||
from nanobot.providers.base import LLMResponse
|
||||
from nanobot.bus.events import InboundMessage
|
||||
|
||||
# Setup
|
||||
bus = MessageBus()
|
||||
sessions = SessionManager(tmp_path)
|
||||
|
||||
# Mock provider
|
||||
mock_provider = Mock()
|
||||
mock_provider.default_model = "mock-model"
|
||||
mock_provider.thinking_budget = 0
|
||||
|
||||
# Mock successful response
|
||||
mock_response = LLMResponse(
|
||||
content="Test response",
|
||||
tool_calls=[],
|
||||
reasoning_content=None
|
||||
)
|
||||
mock_provider.chat = AsyncMock(return_value=mock_response)
|
||||
|
||||
# Create agent loop
|
||||
loop = AgentLoop(
|
||||
provider=mock_provider,
|
||||
bus=bus,
|
||||
session_manager=sessions,
|
||||
workspace=tmp_path
|
||||
)
|
||||
|
||||
# Process message with suppress_output=True
|
||||
msg = InboundMessage(
|
||||
channel="test",
|
||||
sender_id="user",
|
||||
chat_id="123",
|
||||
content="Test message",
|
||||
metadata={"suppress_output": True}
|
||||
)
|
||||
|
||||
response = await loop._process_message(msg)
|
||||
|
||||
# Verify response has suppressed metadata
|
||||
assert response.metadata.get("suppressed") is True
|
||||
|
||||
# Verify session contains signed marker
|
||||
session = sessions.get_or_create("test:123")
|
||||
assistant_messages = [m for m in session.messages if m.get("role") == "assistant"]
|
||||
assert len(assistant_messages) > 0
|
||||
|
||||
last_msg = assistant_messages[-1]["content"]
|
||||
assert last_msg.startswith("[HIDDEN:")
|
||||
|
||||
# Verify signature is valid
|
||||
is_valid, clean = verify_signature(last_msg)
|
||||
assert is_valid is True
|
||||
assert clean == "Test response"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_no_marker_accumulation_with_real_provider(tmp_path):
|
||||
"""
|
||||
CRITICAL TEST: Verify markers don't accumulate when model sees them in context.
|
||||
|
||||
This test uses a semi-realistic provider that sees the context and could
|
||||
potentially copy markers, unlike pure mocks that don't see context at all.
|
||||
"""
|
||||
from unittest.mock import AsyncMock, Mock
|
||||
from nanobot.agent.loop import AgentLoop
|
||||
from nanobot.bus.queue import MessageBus
|
||||
from nanobot.session.manager import SessionManager
|
||||
from nanobot.providers.base import LLMResponse
|
||||
from nanobot.bus.events import InboundMessage
|
||||
import re
|
||||
|
||||
# Setup
|
||||
bus = MessageBus()
|
||||
sessions = SessionManager(tmp_path)
|
||||
|
||||
# Create a provider that SEES context and simulates potential copying behavior
|
||||
class ContextAwareProvider:
|
||||
"""Provider that sees context and could copy markers (simulating real LLM)."""
|
||||
default_model = "test-model"
|
||||
thinking_budget = 0
|
||||
|
||||
def __init__(self):
|
||||
self.call_count = 0
|
||||
self.last_context = None
|
||||
|
||||
def get_default_model(self) -> str:
|
||||
"""Get the default model."""
|
||||
return self.default_model
|
||||
|
||||
async def chat(self, messages, **kwargs):
|
||||
self.call_count += 1
|
||||
self.last_context = messages
|
||||
|
||||
# Count markers in non-system messages (system prompt has 2 mentions in docs)
|
||||
marker_count = sum(
|
||||
msg.get("content", "").count("[HIDDEN:")
|
||||
for msg in messages
|
||||
if isinstance(msg.get("content"), str) and msg.get("role") != "system"
|
||||
)
|
||||
|
||||
# Simulate model behavior: on first call (msg 2), sees 1 marker from msg 1
|
||||
# The model should NOT copy it
|
||||
if self.call_count == 2:
|
||||
# Verify context has exactly 1 marker in assistant messages (from message 1)
|
||||
assert marker_count == 1, f"Expected 1 marker in context, found {marker_count}"
|
||||
|
||||
# Always return clean response (good model behavior)
|
||||
return LLMResponse(
|
||||
content=f"Response {self.call_count}",
|
||||
tool_calls=[],
|
||||
reasoning_content=None
|
||||
)
|
||||
|
||||
provider = ContextAwareProvider()
|
||||
|
||||
# Create agent loop
|
||||
loop = AgentLoop(
|
||||
provider=provider,
|
||||
bus=bus,
|
||||
session_manager=sessions,
|
||||
workspace=tmp_path
|
||||
)
|
||||
|
||||
# Use unique chat_id for this test to avoid pollution from previous runs
|
||||
import time
|
||||
test_chat_id = f"test_accumulation_{int(time.time()*1000)}"
|
||||
|
||||
# Message 1: suppress_output=True → should add signed marker
|
||||
msg1 = InboundMessage(
|
||||
channel="test",
|
||||
sender_id="user",
|
||||
chat_id=test_chat_id,
|
||||
content="Hidden message 1",
|
||||
metadata={"suppress_output": True}
|
||||
)
|
||||
|
||||
await loop._process_message(msg1)
|
||||
|
||||
# Verify message 1 has signed marker
|
||||
session = sessions.get_or_create(f"test:{test_chat_id}")
|
||||
assistant_msgs = [m for m in session.messages if m.get("role") == "assistant"]
|
||||
assert len(assistant_msgs) == 1
|
||||
msg1_content = assistant_msgs[0]["content"]
|
||||
assert msg1_content.startswith("[HIDDEN:")
|
||||
is_valid, clean = verify_signature(msg1_content)
|
||||
assert is_valid is True
|
||||
assert clean == "Response 1"
|
||||
|
||||
# Count markers in session after message 1
|
||||
marker_count_1 = sum(m.get("content", "").count("[HIDDEN:") for m in session.messages if isinstance(m.get("content"), str))
|
||||
assert marker_count_1 == 1, f"Expected 1 marker after msg1, found {marker_count_1}"
|
||||
|
||||
# Message 2: Normal message (context includes message 1 with marker)
|
||||
msg2 = InboundMessage(
|
||||
channel="test",
|
||||
sender_id="user",
|
||||
chat_id=test_chat_id,
|
||||
content="Normal message 2",
|
||||
metadata={}
|
||||
)
|
||||
|
||||
await loop._process_message(msg2)
|
||||
|
||||
# Verify message 2 response does NOT start with [HIDDEN: (model didn't copy)
|
||||
session = sessions.get_or_create(f"test:{test_chat_id}")
|
||||
assistant_msgs = [m for m in session.messages if m.get("role") == "assistant"]
|
||||
assert len(assistant_msgs) == 2
|
||||
msg2_content = assistant_msgs[1]["content"]
|
||||
assert not msg2_content.startswith("[HIDDEN:"), f"Message 2 should not start with [HIDDEN:, got: {msg2_content}"
|
||||
|
||||
# Verify still only 1 marker in session (no accumulation)
|
||||
marker_count_2 = sum(m.get("content", "").count("[HIDDEN:") for m in session.messages if isinstance(m.get("content"), str))
|
||||
assert marker_count_2 == 1, f"Expected 1 marker after msg2, found {marker_count_2} (ACCUMULATION DETECTED)"
|
||||
|
||||
# Message 3: Another suppress_output=True → should add SECOND signed marker
|
||||
msg3 = InboundMessage(
|
||||
channel="test",
|
||||
sender_id="user",
|
||||
chat_id=test_chat_id,
|
||||
content="Hidden message 3",
|
||||
metadata={"suppress_output": True}
|
||||
)
|
||||
|
||||
await loop._process_message(msg3)
|
||||
|
||||
# Verify message 3 has signed marker
|
||||
session = sessions.get_or_create(f"test:{test_chat_id}")
|
||||
assistant_msgs = [m for m in session.messages if m.get("role") == "assistant"]
|
||||
assert len(assistant_msgs) == 3
|
||||
msg3_content = assistant_msgs[2]["content"]
|
||||
assert msg3_content.startswith("[HIDDEN:")
|
||||
is_valid, clean = verify_signature(msg3_content)
|
||||
assert is_valid is True
|
||||
assert clean == "Response 3"
|
||||
|
||||
# Verify exactly 2 markers in session (one from msg1, one from msg3)
|
||||
marker_count_3 = sum(m.get("content", "").count("[HIDDEN:") for m in session.messages if isinstance(m.get("content"), str))
|
||||
assert marker_count_3 == 2, f"Expected 2 markers after msg3, found {marker_count_3}"
|
||||
|
||||
# CRITICAL: Verify no double/triple markers like "[HIDDEN: [HIDDEN: [HIDDEN: message"
|
||||
for msg in session.messages:
|
||||
content = msg.get("content", "")
|
||||
if isinstance(content, str) and "[HIDDEN:" in content:
|
||||
# Count occurrences of [HIDDEN: pattern in this single message
|
||||
hidden_count = content.count("[HIDDEN:")
|
||||
assert hidden_count == 1, f"Message has {hidden_count} [HIDDEN: markers (accumulation): {content[:100]}"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_forged_marker_triggers_rejection(tmp_path):
|
||||
"""Test that forged markers trigger rejection and retry."""
|
||||
from unittest.mock import AsyncMock, Mock
|
||||
from nanobot.agent.loop import AgentLoop
|
||||
from nanobot.bus.queue import MessageBus
|
||||
from nanobot.session.manager import SessionManager
|
||||
from nanobot.providers.base import LLMResponse
|
||||
from nanobot.bus.events import InboundMessage
|
||||
|
||||
# Setup
|
||||
bus = MessageBus()
|
||||
sessions = SessionManager(tmp_path)
|
||||
|
||||
# Mock provider
|
||||
mock_provider = Mock()
|
||||
mock_provider.default_model = "mock-model"
|
||||
mock_provider.thinking_budget = 0
|
||||
|
||||
# First response: model tries to forge marker
|
||||
forged_response = LLMResponse(
|
||||
content="[HIDDEN:deadbeef] Forged message",
|
||||
tool_calls=[],
|
||||
reasoning_content=None
|
||||
)
|
||||
|
||||
# Second response: clean response after correction
|
||||
clean_response = LLMResponse(
|
||||
content="Clean message",
|
||||
tool_calls=[],
|
||||
reasoning_content=None
|
||||
)
|
||||
|
||||
mock_provider.chat = AsyncMock(side_effect=[forged_response, clean_response])
|
||||
|
||||
# Create agent loop
|
||||
loop = AgentLoop(
|
||||
provider=mock_provider,
|
||||
bus=bus,
|
||||
session_manager=sessions,
|
||||
workspace=tmp_path
|
||||
)
|
||||
|
||||
# Process message with suppress_output=True
|
||||
msg = InboundMessage(
|
||||
channel="test",
|
||||
sender_id="user",
|
||||
chat_id="123",
|
||||
content="Test message",
|
||||
metadata={"suppress_output": True}
|
||||
)
|
||||
|
||||
response = await loop._process_message(msg)
|
||||
|
||||
# Verify provider.chat was called twice (initial + retry)
|
||||
assert mock_provider.chat.call_count == 2
|
||||
|
||||
# Verify second call included correction message
|
||||
second_call_messages = mock_provider.chat.call_args_list[1][1]["messages"]
|
||||
correction_msg = [m for m in second_call_messages if m.get("role") == "user" and "rejected" in m.get("content", "").lower()]
|
||||
assert len(correction_msg) > 0
|
||||
|
||||
# Verify final response uses clean content (not forged)
|
||||
session = sessions.get_or_create("test:123")
|
||||
assistant_messages = [m for m in session.messages if m.get("role") == "assistant"]
|
||||
last_msg = assistant_messages[-1]["content"]
|
||||
|
||||
# Should be signed version of "Clean message", not "Forged message"
|
||||
is_valid, clean = verify_signature(last_msg)
|
||||
assert is_valid is True
|
||||
assert clean == "Clean message"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_system_message_handler_uses_signed_markers(tmp_path):
|
||||
"""Test that _process_system_message uses signed markers in suppress mode."""
|
||||
from unittest.mock import AsyncMock, Mock
|
||||
from nanobot.agent.loop import AgentLoop
|
||||
from nanobot.bus.queue import MessageBus
|
||||
from nanobot.session.manager import SessionManager
|
||||
from nanobot.providers.base import LLMResponse
|
||||
from nanobot.bus.events import InboundMessage
|
||||
|
||||
# Setup
|
||||
bus = MessageBus()
|
||||
sessions = SessionManager(tmp_path)
|
||||
|
||||
# Mock provider
|
||||
mock_provider = Mock()
|
||||
mock_provider.default_model = "mock-model"
|
||||
mock_provider.thinking_budget = 0
|
||||
mock_response = LLMResponse(
|
||||
content="System response",
|
||||
tool_calls=[],
|
||||
reasoning_content=None
|
||||
)
|
||||
mock_provider.chat = AsyncMock(return_value=mock_response)
|
||||
|
||||
# Create agent loop
|
||||
loop = AgentLoop(
|
||||
provider=mock_provider,
|
||||
bus=bus,
|
||||
session_manager=sessions,
|
||||
workspace=tmp_path
|
||||
)
|
||||
|
||||
# Process system message with suppress_output=True
|
||||
msg = InboundMessage(
|
||||
channel="system",
|
||||
sender_id="subagent",
|
||||
chat_id="test:123",
|
||||
content="[Subagent completed] Result: OK",
|
||||
metadata={"suppress_output": True}
|
||||
)
|
||||
|
||||
response = await loop._process_system_message(msg)
|
||||
|
||||
# Verify response has suppressed metadata
|
||||
assert response.metadata.get("suppressed") is True
|
||||
|
||||
# Verify session contains signed marker
|
||||
session = sessions.get_or_create("test:123")
|
||||
assistant_messages = [m for m in session.messages if m.get("role") == "assistant"]
|
||||
assert len(assistant_messages) > 0
|
||||
|
||||
last_msg = assistant_messages[-1]["content"]
|
||||
assert last_msg.startswith("[HIDDEN:")
|
||||
|
||||
# Verify signature is valid
|
||||
is_valid, clean = verify_signature(last_msg)
|
||||
assert is_valid is True
|
||||
assert clean == "System response"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_system_message_handler_rejects_forged_markers(tmp_path):
|
||||
"""Test that _process_system_message rejects forged markers."""
|
||||
from unittest.mock import AsyncMock, Mock
|
||||
from nanobot.agent.loop import AgentLoop
|
||||
from nanobot.bus.queue import MessageBus
|
||||
from nanobot.session.manager import SessionManager
|
||||
from nanobot.providers.base import LLMResponse
|
||||
from nanobot.bus.events import InboundMessage
|
||||
|
||||
# Setup
|
||||
bus = MessageBus()
|
||||
sessions = SessionManager(tmp_path)
|
||||
|
||||
# Mock provider
|
||||
mock_provider = Mock()
|
||||
mock_provider.default_model = "mock-model"
|
||||
mock_provider.thinking_budget = 0
|
||||
|
||||
# First response: model tries to forge marker
|
||||
forged_response = LLMResponse(
|
||||
content="[HIDDEN:deadbeef] Forged system message",
|
||||
tool_calls=[],
|
||||
reasoning_content=None
|
||||
)
|
||||
|
||||
# Second response: clean response after correction
|
||||
clean_response = LLMResponse(
|
||||
content="Clean system response",
|
||||
tool_calls=[],
|
||||
reasoning_content=None
|
||||
)
|
||||
|
||||
mock_provider.chat = AsyncMock(side_effect=[forged_response, clean_response])
|
||||
|
||||
# Create agent loop
|
||||
loop = AgentLoop(
|
||||
provider=mock_provider,
|
||||
bus=bus,
|
||||
session_manager=sessions,
|
||||
workspace=tmp_path
|
||||
)
|
||||
|
||||
# Process system message with suppress_output=True
|
||||
msg = InboundMessage(
|
||||
channel="system",
|
||||
sender_id="subagent",
|
||||
chat_id="test:456",
|
||||
content="[Subagent completed] Result: OK",
|
||||
metadata={"suppress_output": True}
|
||||
)
|
||||
|
||||
response = await loop._process_system_message(msg)
|
||||
|
||||
# Verify provider.chat was called twice (initial + retry)
|
||||
assert mock_provider.chat.call_count == 2
|
||||
|
||||
# Verify second call included correction message
|
||||
second_call_messages = mock_provider.chat.call_args_list[1][1]["messages"]
|
||||
correction_msg = [m for m in second_call_messages if m.get("role") == "user" and "rejected" in m.get("content", "").lower()]
|
||||
assert len(correction_msg) > 0
|
||||
|
||||
# Verify final response uses clean content (not forged)
|
||||
session = sessions.get_or_create("test:456")
|
||||
assistant_messages = [m for m in session.messages if m.get("role") == "assistant"]
|
||||
last_msg = assistant_messages[-1]["content"]
|
||||
|
||||
# Should be signed version of "Clean system response", not "Forged system message"
|
||||
is_valid, clean = verify_signature(last_msg)
|
||||
assert is_valid is True
|
||||
assert clean == "Clean system response"
|
||||
|
||||
Reference in New Issue
Block a user