Add new environments and enhance tool context functionality

- Introduced new environments: Terminal Test Environment and SWE Environment, each with default configurations for testing and software engineering tasks. - Added TerminalBench 2.0 evaluation environment with comprehensive setup for agentic LLMs, including task execution and verification. - Enhanced ToolContext with methods for uploading and downloading files, ensuring binary-safe operations. - Updated documentation across environments to reflect new features and usage instructions. - Refactored existing environment configurations for consistency and clarity.
2026-06-17 09:41:58 +00:00 · 2026-02-10 19:39:05 +00:00 · 2026-02-10 19:39:05 +00:00 · 35ad3146a8
commit 35ad3146a8
parent e8343f2d87
18 changed files with 1428 additions and 19 deletions
--- a/environments/init.py
+++ b/environments/init.py
@ -4,15 +4,18 @@ Hermes-Agent Atropos Environments
 Provides a layered integration between hermes-agent's tool-calling capabilities
 and the Atropos RL training framework.

-Layers:
+Core layers:
    - agent_loop: Reusable multi-turn agent loop with standard OpenAI-spec tool calling
    - tool_context: Per-rollout tool access handle for reward/verification functions
    - hermes_base_env: Abstract base environment (BaseEnv subclass) for Atropos
    - tool_call_parsers: Client-side tool call parser registry for Phase 2 (VLLM /generate)

 Concrete environments:
-    - terminal_test_env: Simple file-creation tasks for testing the stack
-    - hermes_swe_env: SWE-bench style tasks with Modal sandboxes
+    - terminal_test_env/: Simple file-creation tasks for testing the stack
+    - hermes_swe_env/: SWE-bench style tasks with Modal sandboxes
+
+Benchmarks (eval-only):
+    - benchmarks/terminalbench_2/: Terminal-Bench 2.0 evaluation
 """

 from environments.agent_loop import AgentResult, HermesAgentLoop