測試 mod
為 Claude Code mod 編寫自動化測試,該測試會觸發事件、存根 Claude Code 的答案並按下按鈕,無需工作階段、登入或網路。
您可以為 mod 編寫自動化測試,並使用 claude plugin test 從您的 shell 執行它們。測試會觸發您的 hooks 處理的事件,並檢查 hooks 執行了什麼,以便在問題到達工作階段之前捕捉它。第一個範例測試來自 Create a mod 的 mod。
編寫測試
測試會載入您的 mod,透過其 hooks 傳送事件,就像 Claude Code 會做的那樣,並檢查 hooks 執行了什麼,無需工作階段、登入或網路。您可以使用 claude plugin test 從 shell 執行測試,每個測試檔案都會匯入測試套件,這是 claude-code/testing 模組中的測試庫。
給每個測試檔案一個以 .test.ts 結尾的名稱,例如 first-mod.test.ts,並將其保存在 plugin 目錄中的任何位置。每個測試檔案至少需要一個 test(),否則執行會失敗,並顯示 declares no test(): nothing ran。測試檔案可以匯入您的 mod 自己的檔案和同級 .ts 幫助程式,因此您可以對純函數進行單元測試,例如遊戲的規則,而無需使用套件。
此測試會觸發兩個工具呼叫,執行來自 Create a mod 的 /tally 命令,並檢查回覆是否計算了兩者。其第一行是一個 stub,它代替 Claude Code 回答工具呼叫。將其保存為 first-mod/tests/first-mod.test.ts:
import { expect, test } from 'claude-code/testing'
test('/tally reports the tool calls the mod has seen', async ($, on) => {
// Answer each tool call in Claude Code's place, so no tool runs
on('tool.call', () => ({ result: 'ok' }))
// Raise two tool calls, which the mod's tool.call hook counts
await $.tool.call({ tool: 'Bash', command: 'ls' })
await $.tool.call({ tool: 'Read', file_path: 'README.md' })
// Run /tally and check the text its hook returns
const answer = await $.command.run({ command: 'tally', args: '' })
expect(answer.text).toBe('Claude has made 2 tool calls since this mod loaded')
})
在您的 shell 中,從 first-mod 目錄執行測試:
claude plugin test
輸出會列出每個測試及其是否通過,以及每次執行都會變化的計時:
tests/first-mod.test.ts:
(pass) /tally reports the tool calls the mod has seen [22.87ms]
1 pass
0 fail
Ran 1 test across 1 file. [0.19s]
每個 $.tool.call 都經過了 mod 的 tool.call hook,該 hook 將其計數加一,並將呼叫傳遞給 stub。沒有 ls 執行,也沒有檔案被讀取。$.command.run 隨後進入 mod 的 command.run hook,answer 是該 hook 返回的物件。
當測試失敗時,命令會以狀態 1 退出,因此它在 CI 中有效。如果您自己的 mods 無法在執行它的 shell 中載入,它會列印一行以 claude plugin test: hooks modules are turned off 開頭的行,並顯示原因,然後以狀態 1 退出。
存根 Claude Code 會回答的內容
在測試中沒有模型、存儲或工具執行,因此無論您的 mod 期望 Claude Code 回答的地方,測試都會使用 stub 提供答案。測試函數為此接收兩個引數:
$:測試自己的$,它代替 Claude Code。它不是 hook 接收的 mods API。它的每個方法都會觸發相同名稱的事件,透過您的 mod 的 hooks 傳送它,並解析為結果:$.tool.call({ tool: 'Bash', command: 'ls' })觸發tool.call。$.command.run、$.prompt.submit、$.session.start和$.turn.complete的工作方式相同,$.classic.Stop和其他$.classic方法會觸發 settings hook 事件。測試無法直接觸發 mods API 呼叫,例如ui.close。透過您的 mod 觸發它,例如按下關閉窗格的按鈕。on:呼叫它來註冊 stubs,這些是代替 Claude Code 回答的 hooks。為 mods API 呼叫命名 stub 時不要使用$.,因此註冊為store.get的 stub 會回答您的 mod 的$.store.get。當您的 mod 呼叫$.model.complete或$.store.get時,stub 會提供答案。
此範例存根一個模型呼叫。該 hook 屬於名為 grader 的 mod,並處理一個 /grade 命令,該命令將句子發送到模型並報告回覆是否以 PASS 開頭。該檔案只包含正在測試的 hook,因此 mod 還需要 plugin.json 和 hooks.json,如 Create a mod 中所示。要在工作階段中輸入 /grade,mod 還必須 register the command:
export function register(on) {
on('command.run', { command: 'grade' }, async ($, e) => {
// e.args is the text typed after /grade
const reply = await $.model.complete({
model: 'haiku',
system: 'Grade the sentence. Start your reply with PASS or FAIL.',
prompt: e.args,
})
const passed = reply.isAnswered && reply.text.startsWith('PASS')
return { text: passed ? 'Passed' : 'Try again' }
})
}
此測試存根模型呼叫以檢查 hook 對通過回覆的處理:
import { expect, test } from 'claude-code/testing'
test('a passing grade is reported', async ($, on) => {
// Answer the mod's $.model.complete call with a fixed reply, so no model runs
on('model.complete', () => ({
value: {
isAnswered: true,
text: 'PASS\nNice sentence.',
usage: { input_tokens: 10, output_tokens: 5, cache_read_input_tokens: 0, cache_creation_input_tokens: 0 },
},
}))
// Run /grade, which makes the mod call the model
const answer = await $.command.run({ command: 'grade', args: 'The cat sat on the mat.' })
expect(answer.text).toBe('Passed')
})
測試通過是因為 hook 的 reply 是 value 下的物件,其 text 以 PASS 開頭。要檢查另一個分支,請新增第二個測試,其 stub 返回以 FAIL 開頭的 text,並期望 Try again。
mods API 呼叫的 stub 返回一個具有 value 欄位的物件,該欄位保存呼叫在您的 mod 中解析的內容:{ value: 7 } 使 $.store.get 解析為 7。Claude Code 事件(例如 turn.step 或 tool.call)的 stub 返回該事件自己的結果,例如 { result: 'ok' }。$.session.send 和 $.prompt.fill 也採用其事件的結果,如表所示。Look up what a stub returns 顯示每個常見名稱採用的形式。兩個錯誤意味著 stub 是錯誤的或遺漏的。失敗的測試的輸出包括一個以 the engine reported: 開頭的區塊,每個錯誤都出現在那裡:
returned neither { value } nor { deny }:mods API 呼叫的 stub 返回了一個裸值no implementation for後跟一個名稱:您的 mod 進行了該呼叫,沒有 stub 回答它
該套件還在記憶體中匯出 mocks,為您回答整個命名空間。mock.clock(on) 回答 $.clock,mock.store(on, { count: 7 }) 從以這些項目開始的存儲中回答 $.store,mock.env(on, { CI: 'true' }) 從這些變數中回答 $.env.get。mock.clock 返回一個您的測試可以推進的模擬時鐘,因此計時器測試不會等待。mock.store 不返回任何內容,因此要檢查您的 mod 保存了什麼,請自己編寫兩個 store stubs,如 drawing test 所做的那樣。
遵循測試套件的規則
測試套件有一些自己的規則,違反其中一個會產生新測試作者首先遇到的錯誤:
-
在測試首次呼叫
$之前註冊每個 stub。 在此之後呼叫on會拋出錯誤,例如on("ui.render") after the test first called $。 -
session.start不會自行執行。 每個測試都以您的模組新鮮載入開始,其 hooks 都未被呼叫,因此模組級變數保持其初始值。如果 hook 依賴於session.start設定的內容,請先觸發它:// Answer the event after your hook passes it on with next(e) on('session.start', () => ({ cwd: '/work' })) // Answer the $.command.register call your hook makes on('command.register', () => ({ value: undefined })) // Raise the event, which runs your session.start hook await $.session.start({ surface: 'terminal', isInteractive: true, cwd: '/work' })第二個 stub 回答
session.starthook(例如 tutorial's)進行的$.command.register呼叫。沒有它,該呼叫會以no implementation for command.register拒絕,套件會跳過您的 hook,因此 hook 中呼叫之後的任何內容都不會執行。測試在該點不會失敗。只有當稍後的檢查失敗時,跳過的 hook 才會在the engine reported:下列出。 -
返回
next(e)的 hook 需要 stub 來回答。 例如,當您的ui.renderhook 返回next(e)時,為了在 Claude 閒置時不繪製任何內容,mounting it 會失敗,並顯示no implementation for ui.render。註冊一個返回元素作為純資料的 stub:// Stands for what Claude Code would draw at the site on('ui.render', () => ({ type: 'Text', props: {}, children: ['drawn by Claude Code'] }))註冊 stub 後,掛載成功,每當您的 hook 返回
next(e)時,ui.find({ type: 'Text' })都會返回該元素。 -
turn.step的 stub 是一個非同步生成器,測試會讀取流直到結束以獲得結果:on('turn.step', async function* ($, e) { // Each yield is one piece of the model's streamed reply yield { kind: 'text', index: 0, text: 'ok' } // The return value is the result of the whole request return { turnId: e.turnId, index: e.index, answer: 'ok', toolUses: [], stopReason: 'end_turn', usage: null } }) // Raise one request to the model, which runs your turn.step hook const stream = $.turn.step({ turnId: 't', index: 0, model: 'claude-test', messageCount: 1 }) // Read every piece until the stream says it's done let step = await stream.next() while (step.done !== true) step = await stream.next() const result = step.value當迴圈結束時,
result是 stub 返回的物件,在您的turn.stephook 有機會更改它之後。這裡result.answer是'ok'。 -
使用工具的名稱和引數作為欄位來觸發工具呼叫,例如
await $.tool.call({ tool: 'Bash', command: 'ls' }),並註冊一個返回{ result }的tool.callstub。
查詢 stub 返回的內容
您的 mod 在測試中進行的每個 mods API 呼叫都需要一個 stub 來回答,除了套件自己回答的少數幾個:$.ui.invalidate 和 $.state 呼叫。對於 $.clock 呼叫,使用 mock.clock(on),否則您的 mod 的 $.clock.now() 會失敗,並顯示 no implementation for clock.now。
此表列出了 mods 最常使用的。第一列是您的 mod 進行的呼叫或它使用 next(e) 傳遞的事件。第二列是傳遞給該名稱下的 on 的函數,因此 $.store.get 列變成 on('store.get', ($, e) => ({ value: saved.get(e.key) }))。stub 中的 '...' 標記您要填入的文字:
| 您的 mod 呼叫或傳遞 | Stub |
|---|---|
$.command.register、$.tool.register、$.ui.toast、$.ui.log、$.ui.status、$.ui.close、$.store.set |
() => ({ value: undefined })。對於 ui.toast 和 ui.log,文字是 e.text。 |
$.store.get |
($, e) => ({ value: saved.get(e.key) }) |
$.fs.read |
($, e) => ({ value: e.path.endsWith('notes.md') ? '# Notes' : '' })。e.path 作為絕對路徑到達,因此與 endsWith 比較。 |
$.ui.open |
() => ({ value: { isPlaced: true } }) |
$.ui.ask |
一個 tool.call stub,因為問題作為對 AskUserQuestion 工具的呼叫到達它:($, e) => ({ result: { answers: { [e.questions[0].question]: 'Run it' } } })。如果您的 mod 傳遞其他工具呼叫,請先檢查 e.tool。 |
$.model.complete |
() => ({ value: { isAnswered: true, text: '...', usage } }) |
$.process.run |
($, e) => ({ value: { exitCode: 0, stdout: '...', stderr: '' } })。e.argv 是引數列表,e.init 保存 cwd 和 timeoutMs。 |
| 任何應該失敗的 mods API 呼叫 | () => ({ deny: 'the reason' }),這使呼叫在您的 mod 中拒絕。拋出的 stub 會被跳過。 |
session.start |
() => ({ cwd: '/work' }) |
turn.start |
($, e) => ({ turnId: e.turnId }) |
tool.call |
() => ({ result: '...' }) |
turn.complete |
() => ({ text: '' })。使用 $.turn.complete({ turnId, answer, durationMs, isAborted: false, usage: null }) 觸發它。 |
prompt.submit |
($, e) => ({ text: e.text }) |
prompt.fill |
() => ({ isFilled: true }) |
$.prompt.read |
() => ({ value: { text: '...', cursor: 0 } }) |
$.ui.copy |
() => ({ value: { isCopied: true } }) |
$.session.messages |
() => ({ value: [{ role: 'assistant', text: '...', toolUses: [] }] }) |
$.session.id、$.agent.list |
() => ({ value: 'abc123' })、() => ({ value: [] }) |
session.send |
() => ({ isDelivered: true })。e.to 作為字串到達,即使您的 mod 傳遞了 { sessionId }。 |
session.receive |
($, e) => ({ text: e.text })。使用 $.session.receive({ origin: { kind: 'peer-send-message' }, text }) 觸發它。 |
ui.render |
() => ({ type: 'Text', props: {}, children: ['...'] }) |
expect 有斷言 toBe、toEqual、toMatch、toMatchObject、toContain、toBeDefined、toBeUndefined 和 toThrow,以及任何斷言之前的 .not。
測試計時器
在計時器上執行工作的 mod 需要測試控制的時鐘,因此測試可以向前移動時間而不是等待。const clock = mock.clock(on) 返回一個從 0 開始的模擬時鐘,只有在您的測試移動它時才會移動。要從另一個時間開始,請以毫秒為單位傳遞它,如 mock.clock(on, { now: 5000 })。時鐘有這些方法:
| 方法 | 它做什麼 |
|---|---|
await clock.advance(1000) |
將時間向前移動該毫秒數,並執行每個到期的計時器 |
await clock.set(5000) |
將時間向前移動到該值,如 advance 會做的那樣 |
clock.now() |
返回時間,這是您的 mod 的 $.clock.now() 解析的內容 |
await clock.settle() |
執行已經到期的計時器,例如一系列零延遲 $.clock.after 呼叫,而不移動時間 |
await clock.sleep(2000) |
在 stub 內,使該 stub 只有在測試推進到那麼遠時才回答,這是您模擬緩慢模型或程序的方式 |
此 hook 屬於名為 countdown 的 mod,並處理一個 /countdown 命令,該命令接受秒數,啟動一個一秒的 $.clock.every 計時器,並在零時顯示 toast。與 grader 一樣,該檔案只包含正在測試的 hook,不註冊命令:
export function register(on) {
on('command.run', { command: 'countdown' }, async ($, e) => {
// e.args is the text typed after /countdown
let left = Number(e.args)
const timer = $.clock.every(1000, () => {
left -= 1
if (left === 0) {
timer.cancel()
$.ui.toast('Time is up')
}
})
// Print nothing in the transcript
return {}
})
}
此測試執行 /countdown 3 並移動模擬時鐘,因此它檢查三秒的行為而無需等待三秒:
import { expect, mock, test } from 'claude-code/testing'
test('the countdown ends with a toast', async ($, on) => {
// Answer every $.clock call from a clock the test controls
const clock = mock.clock(on)
// Collect the text of each toast the mod shows
const toasts: string[] = []
on('ui.toast', ($, e) => {
toasts.push(e.text)
return { value: undefined }
})
await $.command.run({ command: 'countdown', args: '3' })
// After two seconds the timer has fired twice, and no toast is due
await clock.advance(2000)
expect(toasts).toEqual([])
// The third second brings the count to zero
await clock.advance(1000)
expect(toasts).toEqual(['Time is up'])
})
第一個 expect 顯示 toast 不會提前出現,第二個顯示它出現一次。每個 advance 在到期的計時器執行後解析,因此下一行的檢查會看到它們的效果。
測試繪圖
測試可以繪製您的 mod 的 render sites 之一,然後按下、輸入並找到它繪製的元素。$.ui.mount 透過您的 mod 的 ui.render hook 繪製該網站,並返回一個具有每個方法的句柄。要在一個測試中涵蓋多個應用程式,請將 surface 設定為要繪製的應用程式。此測試打開來自 Build a pane with tabs 的窗格,切換標籤,按下按鈕,並檢查終端和桌面應用程式中的計數:
import { expect, test } from 'claude-code/testing'
// What Claude Code passes to a ui.render hook for this pane, apart from the app
const PANE = {
plugin: 'hello-tabs',
component: 'Pane',
requestId: 'hello-tabs',
viewport: { columns: 100, rows: 30 },
props: {
title: 'Hello tabs',
isFocused: true,
bodyColumns: 60,
placement: 'inline',
scroll: { offset: 0, bodyRows: 10 },
view: {},
},
} as const
test('the second tab counts presses and saves the count', async ($, on) => {
// Stub $.store with a Map, so the test can read what the mod saved
const saved = new Map<string, unknown>()
on('store.get', ($, e) => ({ value: saved.get(e.key) }))
on('store.set', ($, e) => {
saved.set(e.key, e.value)
return { value: undefined }
})
// Draw the pane once for each app
for (const surface of ['terminal', 'desktop'] as const) {
const ui = await $.ui.mount({ ...PANE, surface })
// Press the buttons by the key the mod gave them
await ui.press({ key: 'tab-two' })
await ui.press({ key: 'more' })
// The second tab's count line is in the drawing
expect(await ui.find({ type: 'Text', text: /^Count: \d+$/ })).toBeDefined()
await ui.unmount()
}
// One press in each app makes two
expect(saved.get('count')).toBe(2)
})
在您的 shell 中,從 hello-tabs 目錄執行 claude plugin test。當兩個應用程式都繪製計數行且 mod 已保存 2 時,測試通過。計數從第一個應用程式延續到第二個應用程式,因為兩個掛載都使用相同的載入模組。
$.ui.mount 返回的句柄有這些方法,它們透過您給它們的 key 來定址元素:
| 方法 | 它做什麼 |
|---|---|
press({ key: 'more' }) |
按下具有該 key 的 Button |
input({ key: 'new-note', text: 'buy milk' }) |
將文字輸入到具有該 key 的 Input 中並按 Enter。新增 kind: 'change' 以輸入而不提交。 |
select({ key: 'size', value: 'large' }) |
在具有該 key 的 Select 中選擇具有該值的選項 |
find({ key: 'more' }) 或 find({ type: 'Text', text: 'Count: 2' }) |
返回第一個匹配的元素作為 { type, props, children },或 undefined。text 可以是字串或正規表達式。 |
unmount() |
移除繪圖 |
每個方法在您的處理程式完成後解析,因此您可以在下一行檢查結果。將 props 設定為 Claude Code 為該網站傳遞的內容。render sites table 列出每個網站的 props,your build 的類型 有它們的類型。
繪圖測試檢查您的 hook 返回的樹以及它對該應用程式是否有效。它不檢查應用程式如何繪製它,因此也在真實工作階段中查看新佈局。
在 `/clear` 之後測試繪圖
每個測試都以每個 $.state 值在其預設值開始,這是 /clear 留下它們的方式。要測試您的 mod 接下來做什麼,請跳過 session.start,使用 source: 'clear' 觸發 classic.SessionStart,並檢查您的 mod 繪製的內容。
此測試檢查來自 Load a saved value again after /clear 的模組。將其新增到 Test a drawing 中的檔案,其中定義了 PANE。該檔案的第一個測試期望按鈕保存計數,如 Save from more than one session 中的按鈕所做的那樣:
test('the saved count comes back after /clear', async ($, on) => {
// The store already holds a count of 7
on('store.get', () => ({ value: 7 }))
// Answer the event after your hook passes it on with next(e)
on('classic.SessionStart', () => ({}))
// Raise the event that fires after /clear, which runs your hook
await $.classic.SessionStart({ source: 'clear' })
const ui = await $.ui.mount({ ...PANE, surface: 'terminal' })
await ui.press({ key: 'tab-two' })
// The pane shows the stored count, not the default of 0
expect(await ui.find({ type: 'Text', text: 'Count: 7' })).toBeDefined()
})
當您的 classic.SessionStart hook 在窗格繪製之前將儲存的 7 複製到 $.state 時,測試通過。如果您的模組中沒有該 hook,窗格會繪製 Count: 0,find 返回 undefined,測試在 toBeDefined 處失敗。
測試判斷其他 mods 的 mod
您的組織在 prependPlugins 中列出的 mod 可以在另一個 mod 載入之前拒絕它。要測試一個,請設定您的 mod 的層級,並給測試第二個 mod 供您的 mod 接受或拒絕:
tier:在測試檔案的頂部呼叫它一次,如tier('prepend'),以將您的 mod 載入為prepend、append或builtin,其在 order mods run in 中的位置。沒有它,您的 mod 會載入為user。plugins:在測試主體之前將test傳遞一個選項物件。其plugins陣列保存您內聯編寫的 mods,每個都有name和register函數。要在user以外的位置載入一個,請將tier新增到它。
此測試檔案首先載入 policy mod from the admin page。它檢查策略 mod 是否拒絕啟動程序的 mod,並接受不啟動的 mod:
import { expect, test, tier } from 'claude-code/testing'
// Load the mod under test ahead of every other mod
tier('prepend')
// A second mod whose code calls $.process.run, which the policy blocks
const runner = {
name: 'runner',
register(on) {
on('tool.call', async ($, e, next) => {
await $.process.run(['ls'])
return { result: 'runner answered' }
})
},
}
// A second mod that calls nothing the policy blocks
const reader = {
name: 'reader',
register(on) {
on('tool.call', async ($, e, next) => {
return { result: 'reader answered' }
})
},
}
test('refuses a mod that starts a process', { plugins: [runner] }, async ($, on) => {
on('tool.call', () => ({ result: 'claude code answered' }))
let message = ''
try {
// The first call on $ loads the mods, so the refusal is thrown here
await $.tool.call({ tool: 'Bash', command: 'ls' })
} catch (error) {
message = error.message
}
expect(message).toBe('runner: refused by acme-guard: Acme policy: mods may not call process.run')
})
test('admits a mod that starts no process', { plugins: [reader] }, async ($, on) => {
on('tool.call', () => ({ result: 'claude code answered' }))
const out = await $.tool.call({ tool: 'Bash', command: 'ls' })
// The answer comes from reader, which shows that it loaded
expect(out).toEqual({ result: 'reader answered' })
})
在您的 shell 中,從 acme-guard 目錄執行 claude plugin test。當策略 mod 如管理頁面所示時,兩個測試都通過。
該套件在測試首次呼叫 $ 時載入每個 mod。當您的 mod 拒絕一個時,該呼叫會拋出,消息會命名被拒絕的 mod、拒絕它的 mod 和您的原因。在第二個測試中,沒有任何內容被拒絕,因此 reader 在到達 stub 之前回答工具呼叫。
後續步驟
- Troubleshoot a mod:找出為什麼 mod 在工作階段中不執行任何操作
- Mods reference:每個事件的輸入和結果,用於編寫 stubs