SpyBara
Go Premium

plugins/mods/test.md 2026-10-01 23:59 UTC to 2026-10-02 22:59 UTC

This page contains 78 additions and 78 deletions.

2026
Thu 1 23:59 Fri 2 22:59

mod をテストする

イベントを発火し、Claude Code の応答をスタブし、ボタンを押す Claude Code mod の自動テストを、セッション、サインイン、ネットワークなしで作成します。

mod の自動テストを作成し、claude plugin test を使ってシェルから実行できます。テストはフックが処理するイベントを発火し、フックが何をしたかを確認するため、問題がセッションに到達する前に発見できます。最初の例では、mod を作成するで作成した mod をテストします。

テストを書く

テストは mod を読み込み、Claude Code と同じようにそのフックにイベントを送り、フックが何をしたかを確認します。セッションもサインインもネットワークも必要ありません。テストはシェルから claude plugin test で実行し、各テストファイルは claude-code/testing モジュールにあるテストライブラリであるテストキットをインポートします。

各テストファイルには first-mod.test.ts のように .test.ts で終わる名前を付け、プラグインディレクトリ内の任意の場所に保存します。すべてのテストファイルには少なくとも 1 つの test() が必要で、ない場合は declares no test(): nothing ran で実行が失敗します。テストファイルは mod 自身のファイルや同じ階層の .ts ヘルパーをインポートできるため、ゲームのルールなどの単純な関数であれば、キットを使わずにユニットテストできます。

このテストは 2 回のツール呼び出しを発生させ、mod を作成するの /tally コマンドを実行し、その返答が両方を数えていることを確認します。最初の行はスタブで、Claude Code の代わりにツール呼び出しに応答します。first-mod/tests/first-mod.test.ts として保存します。

import { expect, test } from 'claude-code/testing'

test('/tally reports the tool calls the mod has seen', async ($, on) => {
  // Answer each tool call in Claude Code's place, so no tool runs
  on('tool.call', () => ({ result: 'ok' }))

  // Fire two tool calls, which the mod's tool.call hook counts
  await $.tool.call({ tool: 'Bash', command: 'ls' })
  await $.tool.call({ tool: 'Read', file_path: 'README.md' })

  // Run /tally and check the text its hook returns
  const answer = await $.command.run({ command: 'tally', args: '' })
  expect(answer.text).toBe('Claude has made 2 tool calls since this mod loaded')
})

シェルで、first-mod ディレクトリからテストを実行します。

claude plugin test

出力には各テストの名前と成功したかどうかが表示され、所要時間は実行ごとに異なります。

tests/first-mod.test.ts:
(pass) /tally reports the tool calls the mod has seen [22.87ms]

 1 pass
 0 fail
Ran 1 test across 1 file. [0.19s]

各 $.tool.call は mod の tool.call フックを通過し、フックはカウントを 1 増やしてから呼び出しをスタブに渡しました。ls は実行されず、ファイルも読み込まれていません。次に $.command.run が mod の command.run フックに送られ、answer はそのフックが返したオブジェクトです。

テストが失敗するとコマンドはステータス 1 で終了するため、CI でも使えます。実行するシェルで自分の mod を読み込めない場合は、claude plugin test: hooks modules are turned off で始まる行に理由を表示し、ステータス 1 で終了します。

Claude Code が返す応答をスタブにする

テストではモデルもストアもツールも実行されないため、mod が Claude Code からの応答を期待する箇所では、テストがスタブで応答を提供します。そのためにテスト関数は 2 つの引数を受け取ります。

  • $: テスト独自の $ で、Claude Code の役割を果たします。フックが受け取る mods API ではありません。各メソッドは同名のイベントを発生させ、それを mod のフックに送り、結果で解決されます。$.tool.call({ tool: 'Bash', command: 'ls' }) は tool.call を発生させます。$.command.run、$.prompt.submit、$.session.start、$.turn.complete も同様に動作し、$.classic.Stop やその他の $.classic メソッドは設定フックのイベントを発生させます。テストから ui.close のような mods API 呼び出しを直接発生させることはできません。たとえばペインを閉じるボタンを押すなど、mod を通じてトリガーします。
  • on: これを呼び出してスタブを登録します。スタブは Claude Code の代わりに応答するフックです。mods API 呼び出し用のスタブには $. を付けずに名前を付けます。つまり store.get として登録したスタブが mod の $.store.get に応答します。mod が $.model.complete や $.store.get を呼び出すと、スタブが応答を提供します。

この例ではモデル呼び出しをスタブにします。このフックは grader という名前の mod に属し、文をモデルに送って返答が PASS で始まるかどうかを報告する /grade コマンドを処理します。このファイルにはテスト対象のフックしか含まれていないため、mod を作成すると同様に、mod には plugin.json と hooks.json も必要です。セッションで /grade と入力するには、mod でコマンドを登録する必要もあります。

export function register(on) {
  on('command.run', { command: 'grade' }, async ($, e) => {
    // e.args is the text typed after /grade
    const reply = await $.model.complete({
      model: 'haiku',
      system: 'Grade the sentence. Start your reply with PASS or FAIL.',
      prompt: e.args,
    })
    const passed = reply.isAnswered && reply.text.startsWith('PASS')
    return { text: passed ? 'Passed' : 'Try again' }
  })
}

このテストはモデル呼び出しをスタブにして、合格の返答に対してフックが何をするかを確認します。

import { expect, test } from 'claude-code/testing'

test('a passing grade is reported', async ($, on) => {
  // Answer the mod's $.model.complete call with a fixed reply, so no model runs
  on('model.complete', () => ({
    value: {
      isAnswered: true,
      text: 'PASS\nNice sentence.',
      usage: { input_tokens: 10, output_tokens: 5, cache_read_input_tokens: 0, cache_creation_input_tokens: 0 },
    },
  }))

  // Run /grade, which makes the mod call the model
  const answer = await $.command.run({ command: 'grade', args: 'The cat sat on the mat.' })
  expect(answer.text).toBe('Passed')
})

フックの reply は value の下にあるオブジェクトで、その text が PASS で始まるため、テストは成功します。もう一方の分岐を確認するには、スタブが FAIL で始まる text を返す 2 つ目のテストを追加し、Try again を期待値にします。

mods API 呼び出し用のスタブは value フィールドを持つオブジェクトを返し、このフィールドには mod 内でその呼び出しが解決される値を入れます。{ value: 7 } とすると $.store.get は 7 に解決されます。turn.step や tool.call など Claude Code のイベント用のスタブは、{ result: 'ok' } のようにそのイベント自体の結果を返します。表に示すとおり、$.session.send と $.prompt.fill もそれぞれのイベントの結果を受け取ります。よく使われる名前がそれぞれどちらの形式をとるかは、スタブが返すものを調べるを参照してください。次のエラーは、スタブが誤っているか欠けていることを意味します。失敗したテストの出力には the engine reported: という見出しのブロックが含まれ、各エラーはそこに表示されます。

  • returned neither { value } nor { deny }: mods API 呼び出し用のスタブが値をそのまま返した
  • no implementation for の後に名前が続く: mod がその呼び出しを行ったが、応答するスタブがない

キットは、名前空間全体に応答するインメモリのモックもエクスポートしています。mock.clock(on) は $.clock に応答し、mock.store(on, { count: 7 }) は指定したエントリで始まるストアから $.store に応答し、mock.env(on, { CI: 'true' }) は指定した変数から $.env.get に応答します。mock.clock はテストが進めるモッククロックを返すため、タイマーのテストで待つ必要がありません。mock.store は何も返さないため、mod が何を保存したかを確認するには、描画のテストのように 2 つの store スタブを自分で書きます。

テストキットのルールに従う

テストキットには独自のルールがいくつかあり、それらに違反すると、テストを初めて書く人が最初に遭遇するエラーが発生します。

  • テストで最初に $ を呼び出す前に、すべてのスタブを登録します。 その後に on を呼び出すと、on("ui.render") after the test first called $ のようなエラーがスローされます。

  • session.start は自動では実行されません。 各テストは、モジュールが新たに読み込まれ、どのフックもまだ呼び出されていない状態で始まるため、モジュールレベルの変数は初期値を保持しています。フックが session.start で設定される内容に依存している場合は、先にそれを発生させます。

    // Answer the event after your hook passes it on with next(e)
    on('session.start', () => ({ cwd: '/work' }))
    // Answer the $.command.register call your hook makes
    on('command.register', () => ({ value: undefined }))
    // Fire the event, which runs your session.start hook
    await $.session.start({ surface: 'terminal', isInteractive: true, cwd: '/work' })
    

    2 つ目のスタブは、チュートリアルのもののような session.start フックが行う $.command.register 呼び出しに応答します。これがないと、その呼び出しは no implementation for command.register で拒否され、キットはフックをスキップするため、フック内のその呼び出し以降は何も実行されません。その時点ではテストは失敗しません。スキップされたフックは、後のチェックが失敗した場合にのみ the engine reported: の下に表示されます。

  • next(e) を返すフックには、応答するスタブが必要です。 たとえば Claude がアイドル状態のときに何も描画しないために ui.render フックが next(e) を返すと、それをマウントする処理は no implementation for ui.render で失敗します。要素をプレーンなデータとして返すスタブを登録します。

    // Stands for what Claude Code would draw at the site
    on('ui.render', () => ({ type: 'Text', props: {}, children: ['drawn by Claude Code'] }))
    

    スタブを登録するとマウントは成功し、フックが next(e) を返した場合は常に ui.find({ type: 'Text' }) がその要素を返します。

  • turn.step 用のスタブは非同期ジェネレーターであり、テストは結果を得るためにストリームを最後まで読み取ります。

    on('turn.step', async function* ($, e) {
      // Each yield is one piece of the model's streamed reply
      yield { kind: 'text', index: 0, text: 'ok' }
      // The return value is the result of the whole request
      return { turnId: e.turnId, index: e.index, answer: 'ok', toolUses: [], stopReason: 'end_turn', usage: null }
    })
    
    // Fire one request to the model, which runs your turn.step hook
    const stream = $.turn.step({ turnId: 't', index: 0, model: 'claude-test', messageCount: 1 })
    // Read every piece until the stream says it's done
    let step = await stream.next()
    while (step.done !== true) step = await stream.next()
    const result = step.value
    

    ループが終了すると、result はスタブが返したオブジェクトになります。ただし、turn.step フックがそれを変更する機会を経た後のものです。ここでは result.answer は 'ok' です。

  • ツール呼び出しは、ツール名と引数をフィールドとして指定して発生させます。 たとえば await $.tool.call({ tool: 'Bash', command: 'ls' }) のようにし、{ result } を返す tool.call スタブを登録します。

スタブが返す値を調べる

テスト内で mod が行うすべての mods API 呼び出しには、Claude Code の代わりに応答するスタブが必要です。ただし、キット自身が応答する少数の呼び出し、つまり $.ui.invalidate と $.state の呼び出しは例外です。$.clock の呼び出しには mock.clock(on) を使用してください。そうしないと、mod の $.clock.now() が no implementation for clock.now で失敗します。

この表は、mod で最もよく使われるものを示しています。1 列目は、mod が行う呼び出し、または next(e) で渡すイベントです。2 列目は、その名前で on に渡す関数です。たとえば $.store.get の行は on('store.get', ($, e) => ({ value: saved.get(e.key) })) になります。スタブ内の '...' は、自分で埋めるテキストを示します。

mod が呼び出すもの、または渡すもの スタブ
$.command.register、$.tool.register、$.ui.toast、$.ui.log、$.ui.status、$.ui.close、$.store.set () => ({ value: undefined })。ui.toast と ui.log では、テキストは e.text です。
$.store.get ($, e) => ({ value: saved.get(e.key) })
$.fs.read ($, e) => ({ value: e.path.endsWith('notes.md') ? '# Notes' : '' })。e.path は絶対パスで渡されるため、endsWith で比較します。
$.ui.open () => ({ value: { isPlaced: true } })
$.ui.ask 質問は AskUserQuestion ツールの呼び出しとして届くため、tool.call スタブを使います: ($, e) => ({ result: { answers: { [e.questions[0].question]: 'Run it' } } })。mod が他のツール呼び出しも渡す場合は、先に e.tool を確認します。
$.model.complete () => ({ value: { isAnswered: true, text: '...', usage } })
$.process.run ($, e) => ({ value: { exitCode: 0, stdout: '...', stderr: '' } })。e.argv は引数リストで、e.init には cwd と timeoutMs が含まれます。
失敗させたい任意の mods API 呼び出し () => ({ deny: 'the reason' })。これにより mod 内でその呼び出しが拒否されます。例外をスローするスタブは、代わりにスキップされます。
session.start () => ({ cwd: '/work' })
turn.start ($, e) => ({ turnId: e.turnId })
tool.call () => ({ result: '...' })
turn.complete () => ({ text: '' })。$.turn.complete({ turnId, answer, durationMs, isAborted: false, usage: null }) で発生させます。
prompt.submit ($, e) => ({ text: e.text })
prompt.fill () => ({ isFilled: true })
$.prompt.read () => ({ value: { text: '...', cursor: 0 } })
$.ui.copy () => ({ value: { isCopied: true } })
$.session.messages () => ({ value: [{ role: 'assistant', text: '...', toolUses: [] }] })
$.session.id、$.agent.list () => ({ value: 'abc123' })、() => ({ value: [] })
session.send () => ({ isDelivered: true })。mod が { sessionId } を渡した場合でも、e.to は文字列として渡されます。
session.receive ($, e) => ({ text: e.text })。$.session.receive({ origin: { kind: 'peer-send-message' }, text }) で発生させます。
ui.render () => ({ type: 'Text', props: {}, children: ['...'] })

expect には toBe、toEqual、toMatch、toMatchObject、toContain、toBeDefined、toBeUndefined、toThrow のアサーションがあり、いずれの前にも .not を付けられます。

タイマーをテストする

タイマーで処理を実行する mod には、テストが制御できる時計が必要です。これにより、テストは待つ代わりに時間を進めることができます。const clock = mock.clock(on) はモック時計を返します。この時計は 0 から始まり、テストが進めたときにだけ進みます。別の時刻から始めるには、mock.clock(on, { now: 5000 }) のようにミリ秒単位で渡します。この時計には次のメソッドがあります。

メソッド 動作
await clock.advance(1000) 指定したミリ秒数だけ時刻を進め、期限が来た各タイマーを実行します
await clock.set(5000) advance と同様に、時刻をその値まで進めます
clock.now() 時刻を返します。これは mod の $.clock.now() が解決される値です
await clock.settle() 時刻を進めずに、遅延ゼロの $.clock.after 呼び出しの連鎖など、すでに期限が来ているタイマーを実行します
await clock.sleep(2000) スタブ内で使用すると、テストがその時点まで進めたときにだけそのスタブが応答するようになります。これにより、遅いモデルやプロセスをシミュレートできます

このフックは countdown という名前の mod に属し、/countdown コマンドを処理します。このコマンドは秒数を受け取り、1 秒間隔の $.clock.every タイマーを開始し、ゼロになるとトーストを表示します。grader と同様に、このファイルにはテスト対象のフックだけが含まれ、コマンドは登録しません。

export function register(on) {
  on('command.run', { command: 'countdown' }, async ($, e) => {
    // e.args is the text typed after /countdown
    let left = Number(e.args)
    const timer = $.clock.every(1000, () => {
      left -= 1
      if (left === 0) {
        timer.cancel()
        $.ui.toast('Time is up')
      }
    })
    // Print nothing in the transcript
    return {}
  })
}

このテストは /countdown 3 を実行してモック時計を進めるため、3 秒待たずに 3 秒間の動作を確認できます。

import { expect, mock, test } from 'claude-code/testing'

test('the countdown ends with a toast', async ($, on) => {
  // Answer every $.clock call from a clock the test controls
  const clock = mock.clock(on)
  // Collect the text of each toast the mod shows
  const toasts: string[] = []
  on('ui.toast', ($, e) => {
    toasts.push(e.text)
    return { value: undefined }
  })

  await $.command.run({ command: 'countdown', args: '3' })
  // After two seconds the timer has fired twice, and no toast is due
  await clock.advance(2000)
  expect(toasts).toEqual([])
  // The third second brings the count to zero
  await clock.advance(1000)
  expect(toasts).toEqual(['Time is up'])
})

最初の expect はトーストが早く表示されないことを示し、2 つ目はトーストが 1 回表示されることを示します。各 advance は期限が来たタイマーの実行後に解決されるため、次の行のチェックではその効果を確認できます。

描画をテストする

テストでは、mod の描画箇所のいずれかを描画し、描画された要素を押したり、入力したり、検索したりできます。$.ui.mount は mod の ui.render フックを通じてその描画箇所を描画し、これらの操作ごとのメソッドを持つハンドルを返します。1 つのテストで複数のアプリをカバーするには、surface に描画対象のアプリを設定します。次のテストは、タブ付きのペインを作成するのペインを開き、タブを切り替え、ボタンを押して、ターミナルと Desktop アプリでカウントを確認します。

import { expect, test } from 'claude-code/testing'

// What Claude Code passes to a ui.render hook for this pane, apart from the app
const PANE = {
  plugin: 'hello-tabs',
  component: 'Pane',
  requestId: 'hello-tabs',
  viewport: { columns: 100, rows: 30 },
  props: {
    title: 'Hello tabs',
    isFocused: true,
    bodyColumns: 60,
    placement: 'inline',
    scroll: { offset: 0, bodyRows: 10 },
    view: {},
  },
} as const

test('the second tab counts presses and saves the count', async ($, on) => {
  // Stub $.store with a Map, so the test can read what the mod saved
  const saved = new Map<string, unknown>()
  on('store.get', ($, e) => ({ value: saved.get(e.key) }))
  on('store.set', ($, e) => {
    saved.set(e.key, e.value)
    return { value: undefined }
  })

  // Draw the pane once for each app
  for (const surface of ['terminal', 'desktop'] as const) {
    const ui = await $.ui.mount({ ...PANE, surface })
    // Press the buttons by the key the mod gave them
    await ui.press({ key: 'tab-two' })
    await ui.press({ key: 'more' })
    // The second tab's count line is in the drawing
    expect(await ui.find({ type: 'Text', text: /^Count: \d+$/ })).toBeDefined()
    await ui.unmount()
  }

  // One press in each app makes two
  expect(saved.get('count')).toBe(2)
})

シェルで、hello-tabs ディレクトリから claude plugin test を実行します。両方のアプリがカウント行を描画し、mod が 2 を保存していれば、テストは成功します。両方のマウントが同じ読み込み済みモジュールを使用するため、カウントは最初のアプリから 2 番目のアプリへ引き継がれます。

$.ui.mount が返すハンドルには次のメソッドがあり、指定した key で要素を指定します。

メソッド 動作
press({ key: 'more' }) そのキーを持つ Button を押します
input({ key: 'new-note', text: 'buy milk' }) そのキーを持つ Input にテキストを入力し、Enter キーを押します。送信せずに入力するには kind: 'change' を追加します。
select({ key: 'size', value: 'large' }) そのキーを持つ Select で、その値を持つオプションを選択します
find({ key: 'more' }) または find({ type: 'Text', text: 'Count: 2' }) 最初に一致した要素を { type, props, children } として返すか、undefined を返します。text には文字列または正規表現を指定できます。
unmount() 描画を削除します

各メソッドはハンドラーの処理が完了した後に解決されるため、次の行で結果を確認できます。props には、Claude Code がその描画箇所に渡す値を設定します。描画箇所の表に各描画箇所の props が記載されており、ビルド用の型にその型があります。

描画テストでは、フックが返すツリーと、それがそのアプリに対して有効かどうかを確認します。アプリがそれをどのように表示するかは確認しないため、新しいレイアウトは実際のセッションでも確認してください。

`/clear` 後の描画をテストする

各テストは、すべての $.state の値がデフォルトの状態で開始します。これは /clear 後の状態と同じです。mod がその後に何をするかをテストするには、session.start をスキップし、source: 'clear' を指定して classic.SessionStart を発火させ、mod が描画する内容を確認します。

このテストは、/clear 後に保存した値を再度読み込むのモジュールを確認します。PANE が定義されている描画をテストするのファイルに追加してください。そのファイルの最初のテストは、複数のセッションから保存するのボタンと同様に、ボタンがカウントを保存することを前提としています。

test('the saved count comes back after /clear', async ($, on) => {
  // The store already holds a count of 7
  on('store.get', () => ({ value: 7 }))
  // Answer the event after your hook passes it on with next(e)
  on('classic.SessionStart', () => ({}))

  // Fire the event that follows /clear, which runs your hook
  await $.classic.SessionStart({ source: 'clear' })

  const ui = await $.ui.mount({ ...PANE, surface: 'terminal' })
  await ui.press({ key: 'tab-two' })
  // The pane shows the stored count, not the default of 0
  expect(await ui.find({ type: 'Text', text: 'Count: 7' })).toBeDefined()
})

ペインが描画される前に、classic.SessionStart フックが保存された 7 を $.state にコピーしていれば、テストは成功します。モジュールにそのフックがない場合、ペインは Count: 0 を描画し、find は undefined を返し、テストは toBeDefined で失敗します。

ポリシー mod をテストする

組織が prependPlugins に登録している mod は、別の mod が読み込まれる前にその mod を拒否できます。このような mod をテストするには、自分の mod の tier を設定し、その mod が許可または拒否する対象となる 2 つ目の mod をテストに渡します。

  • tier:テストファイルの先頭で tier('prepend') のように一度呼び出すと、自分の mod を prepend、append、builtin のいずれかとして読み込みます。これは mod の実行順序における位置を表します。指定しない場合、mod は user として読み込まれます。
  • plugins:テスト本体の前に、test にオプションオブジェクトを渡します。その plugins 配列には、インラインで記述した mod を含めます。各 mod には name と register 関数を持たせます。user 以外の位置で読み込むには、その mod に tier を追加します。

このテストファイルは、管理者ページのポリシー mod を最初に読み込みます。そして、ポリシー mod がプロセスを開始する mod を拒否し、プロセスを開始しない mod を許可することを確認します。

import { expect, test, tier } from 'claude-code/testing'

// Load the mod under test ahead of every other mod
tier('prepend')

// A second mod whose code calls $.process.run, which the policy blocks
const runner = {
  name: 'runner',
  register(on) {
    on('tool.call', async ($, e, next) => {
      await $.process.run(['ls'])
      return { result: 'runner answered' }
    })
  },
}

// A second mod that calls nothing the policy blocks
const reader = {
  name: 'reader',
  register(on) {
    on('tool.call', async ($, e, next) => {
      return { result: 'reader answered' }
    })
  },
}

test('refuses a mod that starts a process', { plugins: [runner] }, async ($, on) => {
  on('tool.call', () => ({ result: 'claude code answered' }))
  let message = ''
  try {
    // The first call on $ loads the mods, so the refusal is thrown here
    await $.tool.call({ tool: 'Bash', command: 'ls' })
  } catch (error) {
    message = error.message
  }
  expect(message).toBe('runner: refused by acme-guard: Acme policy: mods may not call process.run')
})

test('admits a mod that starts no process', { plugins: [reader] }, async ($, on) => {
  on('tool.call', () => ({ result: 'claude code answered' }))
  const out = await $.tool.call({ tool: 'Bash', command: 'ls' })
  // The answer comes from reader, which shows that it loaded
  expect(out).toEqual({ result: 'reader answered' })
})

シェルで、acme-guard ディレクトリから claude plugin test を実行します。ポリシー mod が管理者ページに示されているとおりであれば、両方のテストが成功します。

キットは、テストで最初に $ が呼び出された時点ですべての mod を読み込みます。自分の mod がいずれかの mod を拒否すると、その呼び出しは例外をスローし、メッセージには拒否された mod、それを拒否した mod、および指定した理由が示されます。2 つ目のテストでは何も拒否されないため、ツール呼び出しがスタブに到達する前に reader が応答します。

次のステップ