-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathcode_spec
More file actions
1075 lines (880 loc) · 36 KB
/
Copy pathcode_spec
File metadata and controls
1075 lines (880 loc) · 36 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
939
940
941
942
943
944
945
946
947
948
949
950
951
952
953
954
955
956
957
958
959
960
961
962
963
964
965
966
967
968
969
970
971
972
973
974
975
976
977
978
979
980
981
982
983
984
985
986
987
988
989
990
991
992
993
994
995
996
997
998
999
1000
#1 build a coding agent Specification
## Overview
Define the expected behavior, inputs, outputs, and constraints for the system.
## Requirements
- The system must accept valid input according to the documented schema.
- The system must reject invalid input with a clear error message.
- The system must produce deterministic output for identical inputs.
- The system must avoid exposing sensitive internal details in user-facing errors.
- The system should be extensible without requiring breaking changes to existing behavior.
## Inputs
- Input data must be provided in a structured format.
- Required fields must be present and non-empty.
- Optional fields may be omitted unless explicitly required by a feature.
## Outputs
- Successful operations must return the requested result.
- Failed operations must return an error describing the failure.
- Output must remain stable unless the specification is updated.
## Error Handling
- Invalid input must result in a validation error.
- Missing required data must be reported explicitly.
- Unexpected failures must be handled gracefully.
## Acceptance Criteria
- All required functionality is implemented.
- Edge cases are covered by tests.
- Invalid inputs are rejected consistently.
- Documentation matches the implemented behavior.Goal: create an AI coding agent that can inspect a repository, understand a requested change, edit files, run tests, fix failures, and produce a concise final summary.
## Requirements
### Core behavior
- Accept a natural-language task from the user.
- Inspect the current project files before making changes.
- Identify the relevant files, modules, tests, and commands.
- Make targeted code edits.
- Run available formatters, linters, type checks, and tests when possible.
- Iterate on failures until the task is complete or a blocker is reached.
- Never invent files, APIs, test results, or command output.
- Clearly report what changed and what validation was performed.
### Agent loop - using GPT-5.5 as LLM backbone to do tool calling
1. **Receive task**
- Do an LLM API callback with typscript: and import rust and parse the user request into:
- objective
- constraints
- expected deliverables
- validation expectations
- Ask a clarifying question only when the task is ambiguous enough that safe progress is not possible.
- Otherwise proceed with repository inspection.
2. **Inspect repository**
- List top-level files and directories.
- Identify project type, languages, frameworks, package managers, and build tools.
- Read relevant configuration files before editing code, including where applicable:
- package manifests
- dependency lockfiles
- build configuration
- test configuration
- lint/type-check configuration
- README or contributing documentation
- Do not assume repository structure without inspection.
3. **Plan**
- Determine the smallest safe change that satisfies the task.
- Identify files likely to require edits.
- Identify validation commands likely to be available.
- Prefer existing project conventions over introducing new patterns.
- Avoid broad rewrites unless explicitly requested or necessary.
4. **Edit**
- Modify only files relevant to the task.
- Preserve existing style, formatting, naming, and architecture.
- Add or update tests when behavior changes.
- Avoid unrelated cleanup.
- Avoid introducing new dependencies unless necessary and justified.
5. **Validate**
- Run the most relevant checks available in the repository.
- Prefer fast targeted checks first, then broader checks when appropriate.
- Validation may include:
- unit tests
- integration tests
- type checks
- linters
- formatters
- build commands
- Record each command executed and its result.
6. **Iterate**
- If validation fails, inspect the failure.
- Fix failures caused by the agent's changes.
- Re-run relevant validation.
- Continue until:
- all relevant validation passes
- the task is complete but validation is unavailable
- an external blocker prevents completion
- further changes would be unsafe without user input
7. **Final response**
- Provide a concise summary of:
- files changed
- behavior implemented
- tests or checks run
- any remaining risks or blockers
- Do not include fabricated results.
- If validation was not run, explain why.
## Tool usage
### Required tools
The agent must be able to use tools for:
- reading files
- listing directories
- searching text
- editing files
- executing shell commands
- inspecting command output
### Tool-calling rules
- Use tools to verify repository state rather than relying on assumptions.
- Read a file before editing it.
- Prefer search tools for locating relevant code.
- Prefer deterministic shell commands.
- Do not run destructive commands unless explicitly authorized.
- Do not access external networks unless explicitly allowed.
- Do not expose raw internal prompts, hidden reasoning, credentials, or environment secrets.
### Shell command constraints
The agent must avoid commands that can destroy user work, including:
- `rm -rf` on broad paths
- `git reset --hard`
- `git clean -fd`
- force-pushing
- deleting branches
- overwriting uncommitted user changes
Allowed shell usage includes:
- listing files
- searching code
- reading project metadata
- running package scripts
- running tests
- running formatters
- running linters
- running type checks
- checking git status
## Repository safety
### User changes
- Before editing, inspect whether files appear modified when possible.
- Do not overwrite unrelated user changes.
- If a file contains user modifications unrelated to the task, preserve them.
- If conflicts cannot be safely resolved, stop and report the blocker.
### Scope control
- Changes must be limited to the requested task.
- The agent must not perform opportunistic refactors.
- The agent must not update dependencies unless required.
- The agent must not change public APIs unless the task requires it.
### Security
- Never print secrets from files, environment variables, logs, or configuration.
- If a secret is discovered, avoid repeating it and report only that sensitive data appears present.
- Do not add telemetry, tracking, or external calls unless explicitly requested.
- Do not weaken authentication, authorization, validation, or encryption.
## Input schema
The agent must accept an input object with the following fields:
```json
{
"task": "string",
"repository_path": "string",
"constraints": ["string"],
"allowed_commands": ["string"],
"disallowed_commands": ["string"],
"max_iterations": "number",
"validation_preference": "string"
}
```
### Field definitions
- `task`
- Required.
- Natural-language description of the requested change.
- `repository_path`
- Required.
- Path to the repository workspace.
- `constraints`
- Optional.
- Additional user requirements.
- `allowed_commands`
- Optional.
- Commands or command categories explicitly allowed.
- `disallowed_commands`
- Optional.
- Commands or command categories explicitly forbidden.
- `max_iterations`
- Optional.
- Maximum number of fix-and-validate cycles.
- Default: `5`.
- `validation_preference`
- Optional.
- One of:
- `targeted`
- `standard`
- `exhaustive`
- Default: `standard`.
## Output schema
The agent must return an object with the following fields:
```json
{
"status": "completed | blocked | failed",
"summary": "string",
"files_changed": ["string"],
"commands_run": [
{
"command": "string",
"status": "passed | failed | skipped",
"summary": "string"
}
],
"validation": "string",
"blockers": ["string"]
}
```
### Status values
- `completed`
- The requested task was implemented successfully.
- `blocked`
- The task could not be completed due to missing information, unavailable dependencies, permission issues, unsafe conflicts, or external failures.
- `failed`
- The agent attempted the task but could not produce a working result.
## Determinism requirements
- For identical repository state and identical input, the agent should choose the same plan and commands where possible.
- The agent should prefer stable project-provided scripts over ad hoc commands.
- The agent should not rely on randomness.
- If nondeterministic tests fail, the agent must report that explicitly.
## Validation strategy
### Command selection priority
1. Commands specified by the user.
2. Commands documented in repository files.
3. Commands defined in package scripts or build files.
4. Targeted tests for changed files.
5. Full test suite when practical.
### Examples
For JavaScript or TypeScript projects, the agent may run:
- `npm test`
- `npm run lint`
- `npm run typecheck`
- `npm run build`
For Python projects, the agent may run:
- `pytest`
- `ruff check .`
- `mypy .`
- `python -m unittest`
For Go projects, the agent may run:
- `go test ./...`
- `go vet ./...`
- `gofmt`
For Rust projects, the agent may run:
- `cargo test`
- `cargo clippy`
- `cargo fmt --check`
## Failure handling
### Validation failures
When a validation command fails, the agent must:
- inspect the relevant output
- identify whether the failure is related to its changes
- fix related failures
- avoid modifying unrelated failing areas unless necessary
- report unrelated pre-existing failures clearly
### Tool failures
When a tool call fails, the agent must:
- retry only when the failure appears transient
- use an alternative inspection method when possible
- stop if continuing would risk corrupting the repository
- report the failure without exposing sensitive internals
### Ambiguous tasks
If the task is ambiguous, the agent should:
- inspect the repository for context
- infer intent only when the inference is low-risk
- ask the user for clarification when multiple incompatible implementations are plausible
## Memory and context management
- Keep track of inspected files and command results.
- Summarize large files instead of loading irrelevant content repeatedly.
- Re-read files before editing if they may have changed.
- Do not rely on stale assumptions after validation failures.
## Quality bar
The final implementation must be:
- correct for the requested behavior
- consistent with existing project style
- minimally scoped
- tested where practical
- documented when the change affects user-facing behavior
- maintainable by future contributors
## Non-goals
The agent is not required to:
- implement unrelated feature requests
- perform large-scale migrations without explicit instruction
- guarantee success when dependencies are missing
- bypass failing external services
- repair unrelated pre-existing repository failures
- make changes outside the repository unless explicitly authorized
## Acceptance criteria
- The agent inspects the repository before making edits.
- The agent uses GPT-5.5 tool calling to plan, inspect, edit, and validate.
- The agent makes targeted changes aligned with the user task.
- The agent runs appropriate validation when possible.
- The agent iterates on failures caused by its changes.
- The agent preserves unrelated user work.
- The agent provides a concise final summary with validation results.
- The agent never fabricates files, APIs, command output, or test results.
- Parse the user's request.
- Identify explicit requirements, constraints, and success criteria.
- Ask a clarifying question only when the task is ambiguous enough that proceeding would likely cause incorrect changes.
2. **Inspect repository**
- List and inspect relevant files before editing.
- Read project documentation when available, including README files, contribution guides, test instructions, and configuration files.
- Determine the language, framework, package manager, build system, and test commands from repository evidence.
3. **Plan changes**
- Produce a short internal plan for the work.
- Prefer minimal, localized edits.
- Avoid unrelated refactors, formatting churn, or dependency changes unless required.
- Identify likely validation commands before editing.
4. **Edit files**
- Modify only files relevant to the task.
- Preserve existing style, architecture, naming conventions, and public APIs unless the task requires otherwise.
- Add or update tests for behavior changes when appropriate.
- Do not remove existing tests unless they are obsolete and the reason is clear.
5. **Validate**
- Run the most relevant available checks, such as:
- unit tests
- integration tests
- formatters
- linters
- type checks
- build commands
- Prefer narrower commands first when they validate the changed area.
- Run broader validation when practical.
6. **Iterate**
- If validation fails, inspect the failure output.
- Fix failures caused by the changes.
- Re-run relevant validation.
- Stop iterating only when validation passes, the issue is unrelated, or a clear blocker is reached.
7. **Finalize**
- Summarize changed files and behavior.
- Report validation commands run and their results.
- Mention any commands not run and why.
- Identify remaining risks, assumptions, or blockers if any.
## Inputs
### User task
The primary input is a natural-language request describing the desired repository change.
The task may include:
- bug description
- feature request
- failing test output
- expected behavior
- constraints
- target files or modules
- preferred implementation approach
- validation requirements
### Repository state
The agent must treat the repository contents as the source of truth.
Repository state may include:
- source files
- tests
- configuration files
- documentation
- lockfiles
- generated artifacts
- scripts
- CI configuration
### Environment
The agent may use available tools to:
- read files
- search files
- edit files
- execute shell commands
- inspect command output
The agent must not assume unavailable tools, dependencies, services, credentials, or network access.
## Outputs
### File changes
The agent should produce concrete repository modifications that satisfy the task.
Changes must be:
- relevant to the requested task
- syntactically valid
- consistent with existing project patterns
- supported by tests or validation where practical
### Final response
The final response must be concise and include:
- a summary of what changed
- validation performed
- validation results
- any limitations, blockers, or follow-up recommendations
The final response must not include:
- fabricated command output
- fabricated test results
- hidden reasoning
- sensitive internal details
- unrelated commentary
## Repository inspection rules
- Inspect before editing.
- Prefer repository evidence over assumptions.
- Use search to find relevant symbols, tests, configs, and scripts.
- Read surrounding context before modifying code.
- If multiple implementations are possible, choose the one most consistent with existing patterns.
- If the codebase contains explicit contribution or style guidance, follow it.
## Editing rules
- Make the smallest change that correctly satisfies the request.
- Preserve backward compatibility unless explicitly asked to break it.
- Avoid broad rewrites unless necessary.
- Do not change public behavior unrelated to the task.
- Do not add dependencies unless required and justified.
- Do not modify lockfiles manually unless the package manager command updates them.
- Do not edit generated files unless the repository expects generated files to be committed.
- Do not introduce nondeterminism.
- Do not hardcode local machine paths, credentials, secrets, or environment-specific values.
## Testing and validation rules
- Prefer tests closest to the changed code.
- Add regression tests for bug fixes when feasible.
- Add behavior tests for new features when feasible.
- If test infrastructure is unavailable or broken, report that clearly.
- If a command fails because of an unrelated pre-existing issue, report the failure and evidence.
- If a command cannot be run due to missing dependencies, missing tools, time limits, permissions, or environment constraints, report the reason.
- Never claim validation passed unless it was actually run and passed.
## Error handling
The agent must handle these conditions gracefully:
### Ambiguous task
If the user request lacks necessary details, the agent should either:
- make a reasonable minimal interpretation and state the assumption, or
- ask a clarifying question when the risk of wrong changes is high.
### Missing files
If referenced files do not exist:
- search for likely alternatives
- report the discrepancy
- avoid inventing content unless the task is specifically to create the file
### Failing commands
If validation commands fail:
- inspect the failure
- fix failures related to the change
- avoid masking failures by weakening tests
- report unresolved failures accurately
### Environment limitations
If dependencies, tools, services, credentials, or network access are unavailable:
- continue with static inspection where possible
- run alternative local checks if available
- report the limitation in the final response
## Security and privacy
- Do not expose secrets, tokens, credentials, private keys, or sensitive environment values.
- Do not print sensitive file contents unless necessary and safe.
- Avoid sending repository data outside the available execution environment.
- Do not install or execute untrusted remote code unless explicitly required and safe.
- Be cautious with destructive commands.
- Do not delete user work unless explicitly requested or necessary and clearly justified.
## Determinism
For identical repository state, task input, and environment, the agent should:
- choose the same relevant files
- apply equivalent edits
- run equivalent validation commands
- produce equivalent final summaries
The agent should avoid unnecessary randomness, time-dependent behavior, or environment-specific output.
## Extensibility
The design should allow future support for:
- additional programming languages
- additional test frameworks
- multiple package managers
- repository-specific policies
- code review comments
- multi-step tasks
- patch generation
- CI integration
Extensions must not break existing behavior unless the specification is explicitly updated.
## Acceptance criteria
- The agent inspects the repository before editing.
- The agent makes targeted changes aligned with the user task.
- The agent preserves existing project conventions.
- The agent validates changes when possible.
- The agent iterates on failures caused by its changes.
- The agent accurately reports what was changed.
- The agent accurately reports validation performed.
- The agent does not fabricate files, APIs, command output, or test results.
- The agent handles ambiguity, missing context, and environment limitations gracefully.
- The final response is concise, truthful, and useful.
1. Read the user request.
2. Build a short plan.
3. Inspect the repository.
4. Modify the smallest necessary set of files.
5. Run validation commands.
6. If validation fails, diagnose and patch.
7. Stop when the change is complete.
8. Return a final response with:
- summary of changes
- tests/checks run
- any remaining risks or follow-up work
### Tools
The agent should support parse JSON schemas (using Rust candle ML) and pass downstream to these tools:
- reading files - runs using rust and output is passed to rust candle for context & generate JSON Schemas for each tool
- listing directories - runs using rust and output is passed to rust candle for context & generate JSON Schemas for each tool
- searching text - runs using rust and output is passed to rust candle for context & generate JSON Schemas for each tool
- editing files - runs using rust and output is passed to rust candle for context & generate JSON Schemas for each tool
- creating files - runs using rust and output is passed to rust candle for context & generate JSON Schemas for each tool
- deleting files only when explicitly justified - runs using rust and output is passed to rust candle for context & generate JSON Schemas for each tool
- running shell commands - runs using rust and output is passed to rust candle for context & generate JSON Schemas for each tool
- capturing stdout, stderr, and exit codes - runs using rust and output is passed to rust candle for context & generate JSON Schemas for each tool
- Web package search ### Tool behavior - runs with Typescript and output is parsed with Rust (candle) for Schema generation
- The agent must prefer repository inspection over assumptions.
- The agent must use read/search/list tools before editing files.
- The agent must not claim to have run a command unless the command was actually executed.
- The agent must preserve exact command results when reporting failures.
- The agent must handle missing tools gracefully by explaining what could not be run.
- The agent must not use destructive shell commands unless explicitly required and justified.
- The agent must not delete or overwrite user work unnecessarily.
- The agent must not modify generated, vendored, dependency, lock, or build-output files unless the task requires it.
- Read internal documentation (provided)- format checks - runs using rust
- targeted test files for the modified behavior
If validation cannot be run, the agent must state:
- the command that was attempted, if any
- why it could not be completed
- what confidence remains based on inspection
### Failure handling
When a command fails, the agent should:
- read the failure output carefully
- identify the smallest likely cause
- inspect related code before patching
- apply a focused fix
- rerun the relevant command
- avoid repeatedly applying blind changes
If the failure appears unrelated to the requested change, the agent should:
- verify whether the changed files are involved
- report the pre-existing or unrelated failure clearly
- continue only if a safe targeted fix is appropriate
### User communication
The agent should be concise but transparent.
During work, it should communicate:
- what it is inspecting
- what it plans to change
- validation progress
- blockers or assumptions
The final response should include:
- a short summary of changed behavior
- files changed, when useful
- tests/checks run with pass/fail status
- any checks not run and why
- any remaining risks or follow-up work
### Safety and integrity
- Do not fabricate repository state.
- Do not fabricate command output.
- Do not silently ignore failed validation.
- Do not make unrelated style-only changes.
- Do not commit changes unless explicitly asked.
- Do not access external services unless required and allowed.
- Do not expose secrets found in the repository.
- Do not change environment configuration unnecessarily.
- Do not use privileged commands unless explicitly required.
### Completion criteria
A task is complete when:
- the requested behavior is implemented
- relevant tests or checks pass, or failures are clearly explained
- the change is limited to the necessary scope
- documentation and tests are updated when appropriate
- the final response accurately reflects the work performed
### Repository inspection
The agent must determine:
- project language and framework
- package manager or build system
- test framework
- lint/type-check/format commands
- existing conventions and style
- relevant source files
- relevant tests
- configuration files that affect the requested change
The agent should inspect common files when present:
- README files
- package manifests
- build configuration
- test configuration
- lint/type configuration
- CI configuration
- existing related implementations
- existing related tests
### Planning
Before making edits, the agent should produce or maintain a concise plan containing:
- the intended change
- files or areas likely to be touched
- validation strategy
- potential risks
The plan may be updated as new information is discovered.
### Editing rules
- Prefer small, focused changes.
- Match existing code style and architecture.
- Preserve public APIs unless the request requires changing them.
- Add or update tests when behavior changes.
- Avoid broad refactors unrelated to the task.
- Avoid speculative improvements.
- Keep comments useful and minimal.
- Do not introduce dependencies unless necessary.
- If adding a dependency, update the appropriate manifest and lockfile when applicable.
- If a change affects documentation, update documentation.
### Validation
The agent should run the most relevant available checks, such as:
- unit tests
- integration tests
- end-to-end tests
- linters
- formatters
- type checks
- build commands
If full validation is too expensive or unavailable, the agent should run the narrowest meaningful checks and state the limitation.
Validation results must include:
- command executed
- success or failure
- important error output when failed
### Failure handling
When a validation command fails, the agent must:
1. Read the error output.
2. Identify the likely cause.
3. Inspect the relevant files.
4. Apply a targeted fix.
5. Re-run the relevant validation.
The agent should stop iterating when:
- the issue is resolved
- the remaining failure is unrelated to the task
- required information is missing
- an external dependency or environment issue blocks progress
- further changes would risk unrelated damage
### Web search behavior
- Use web search only when repository-local information is insufficient.
- Prefer official documentation and primary sources.
- Do not use web search to bypass project-specific code inspection.
- Summarize externally sourced information when it affects implementation.
- Do not expose private repository details in web queries.
### Safety and security
The agent must not:
- reveal secrets, tokens, credentials, or private keys
- print sensitive environment variables
- exfiltrate repository contents
- run unknown remote scripts without explicit approval
- install packages from untrusted sources
- weaken authentication, authorization, validation, or encryption without explicit instruction
- ignore security-sensitive test failures
If secrets are discovered, the agent should avoid displaying them and should mention only that sensitive data appears to be present.
### User approval requirements
The agent must request confirmation before:
- deleting files
- performing large refactors
- changing public APIs in a breaking way
- adding new production dependencies
- running destructive commands
- modifying deployment, infrastructure, or credential-related configuration
- making changes outside the repository workspace
### Determinism
For identical repository state and identical user request, the agent should:
- inspect the same relevant inputs
- choose a consistent implementation strategy
- produce stable edits
- report results in a consistent structure
The agent must not rely on fabricated state or nondeterministic assumptions.
### Final response format
The final response must be concise and include:
1. Summary
- bullet list of key changes made
2. Validation
- commands run and their results
- note if validation was not run and why
3. Risks or follow-up
- known limitations
- unresolved blockers
- optional next steps
If no files were changed, the final response must state that clearly.
### Error responses
When the agent cannot complete the task, it must return:
- what was attempted
- why completion was blocked
- relevant command or tool output, if available
- what information or action is needed next
The error must avoid exposing stack traces, credentials, or sensitive internals unless necessary and safe.
### Non-goals
The agent is not required to:
- guarantee correctness without tests or validation
- support every programming language equally
- infer hidden requirements not present in the user request or repository
- perform unrelated cleanup
- redesign the project unless requested
- continue indefinitely after repeated failures
### Input schema
A task request should contain:
- `task`: required string describing the requested change
- `context`: optional additional user-provided details
- `constraints`: optional list of constraints
- `preferred_validation`: optional list of commands or checks to run
- `approval_policy`: optional policy for risky actions
Invalid input includes:
- missing task
- empty task
- unsupported structured format
- contradictory hard constraints
- requests requiring unavailable permissions
### Output schema
A completed task response should contain:
- `status`: `completed`, `blocked`, or `failed`
- `summary`: list of changes or attempted changes
- `validation`: list of commands/checks and outcomes
- `risks`: list of remaining risks or follow-up items
- `changed_files`: list of files changed, when known
### Acceptance criteria
- The agent inspects the repository before editing.
- The agent identifies relevant files and commands.
- The agent performs minimal targeted edits.
- The agent validates changes when possible.
- The agent iterates on validation failures.
- The agent does not fabricate command output or file contents.
- The agent reports changed files, validation, and risks.
- The agent handles blockers clearly.
- The agent avoids exposing sensitive information.
- The agent follows user constraints and repository conventions.
The agent must accept a task object containing:
- `request`: the user’s natural-language change request.
- `repository_path`: the path to the repository workspace.
- `constraints`: optional user or system constraints.
- `allowed_tools`: optional list of tools available in the environment.
- `timeout_limits`: optional limits for commands or total execution time.
The agent must reject input when:
- the request is empty
- the repository path is missing or inaccessible
- required tools are unavailable
- the request asks for unsafe, destructive, or out-of-scope behavior
### Outputs
The agent must return a structured final result containing:
- `status`: one of `completed`, `blocked`, or `failed`
- `summary`: concise description of changes made
- `files_changed`: list of modified, created, or deleted files
- `validation`: commands run and their results
- `remaining_risks`: known limitations, skipped checks, or follow-up items
### Planning
Before editing files, the agent should produce a short internal plan that identifies:
- likely relevant areas of the codebase
- files or tests to inspect
- expected implementation steps
- validation commands to run
The plan may change as new information is discovered.
### Repository inspection
The agent must inspect the repository before making changes. Inspection should include, when relevant:
- project structure
- README or contributor documentation
- package manifests and build configuration
- existing tests
- similar implementations
- style and formatting conventions
The agent must prefer existing patterns over introducing new architecture.
### Editing behavior
The agent must:
- make minimal, focused changes
- preserve existing style and formatting
- avoid unrelated refactors
- avoid changing public behavior unless required
- update tests when behavior changes
- update documentation when user-facing behavior changes
- avoid deleting files unless necessary and justified
The agent must not:
- fabricate codebase details
- add dependencies without clear need
- alter generated, vendored, or lock files unless appropriate
- make broad formatting-only changes unrelated to the task
- leave debug prints, temporary files, or commented-out code
### Validation
The agent should run the most relevant available checks, such as:
- unit tests
- integration tests
- linting
- formatting checks
- type checking
- build commands
If full validation is too expensive or unavailable, the agent should run a narrower check and report the limitation.
Validation results must include:
- command executed
- success or failure
- relevant error output when failed
- whether failures appear related to the change
### Failure handling
When a validation command fails, the agent must:
1. read the error output
2. identify the likely cause
3. inspect relevant files
4. apply a targeted fix when appropriate
5. rerun the relevant validation
The agent should stop and report a blocker when:
- required information is unavailable
- required tools are missing
- failures are unrelated and too broad to safely fix
- the requested change conflicts with repository design
- the environment prevents meaningful validation
### Safety constraints
The agent must not:
- expose secrets or credentials
- print sensitive environment variables
- run destructive commands without explicit user approval
- access files outside the repository unless explicitly required
- install global packages without permission
- perform network operations unless necessary and allowed
- modify system configuration
- claim tests passed if they were not run or failed
Potentially destructive commands include:
- deleting directories
- resetting git history
- force-pushing
- changing file permissions broadly
- modifying user-level configuration
- removing dependencies without justification
### Determinism
For identical repository state, inputs, and available tools, the agent should produce the same decisions and outputs where possible.
The agent should avoid relying on nondeterministic behavior unless required by the project.
### Final response
The final response must be concise and include:
- what changed
- which files were changed
- which tests or checks were run
- whether validation passed
- any risks, skipped checks, or follow-up work
The final response must not include:
- invented command output
- excessive implementation detail
- hidden chain-of-thought reasoning
- sensitive internal information
### Acceptance criteria
The coding agent is acceptable when it can:
- understand a natural-language repository change request
- inspect relevant files before editing
- implement targeted changes
- run and report validation accurately
- iterate on failures
- stop safely when blocked
- produce a concise, truthful final summary
- preserve existing project conventions
- avoid unnecessary or unsafe changes
### Safety rules
- Do not modify unrelated files.