SpyBara
Go Premium

plugins/mods/test.md 2026-09-30 23:00 UTC to 2026-10-01 21:59 UTC

This page contains 422 additions and 0 deletions.

2026
Thu 1 23:02

Testare un mod

Scrivi test automatizzati per un mod Claude Code che generano eventi, simulano le risposte di Claude Code e premono pulsanti, senza sessione, accesso o rete.

Puoi scrivere test automatizzati per un mod ed eseguirli dalla tua shell con claude plugin test. Un test genera gli eventi che i tuoi hook gestiscono e controlla cosa hanno fatto gli hook, così puoi individuare un problema prima che raggiunga una sessione. Il primo esempio testa il mod da Create a mod.

Scrivi un test

Un test carica il tuo mod, invia eventi attraverso i suoi hook nel modo in cui Claude Code farebbe, e controlla cosa hanno fatto gli hook, senza una sessione, un accesso o una rete. Esegui i test dalla tua shell con claude plugin test, e ogni file di test importa il test kit, una libreria di test nel modulo claude-code/testing.

Dai a ogni file di test un nome che termina in .test.ts, come first-mod.test.ts, e salvalo ovunque nella directory del plugin. Ogni file di test ha bisogno di almeno un test(), altrimenti l'esecuzione fallisce con declares no test(): nothing ran. Un file di test può importare i tuoi file mod e helper .ts fratelli, così puoi fare unit test di funzioni semplici, come le regole di un gioco, senza il kit.

Questo test genera due chiamate di strumento, esegue il comando /tally da Create a mod, e controlla che la risposta conti entrambe. La sua prima riga è uno stub, che risponde alle chiamate di strumento al posto di Claude Code. Salvalo come first-mod/tests/first-mod.test.ts:

import { expect, test } from 'claude-code/testing'

test('/tally reports the tool calls the mod has seen', async ($, on) => {
  // Answer each tool call in Claude Code's place, so no tool runs
  on('tool.call', () => ({ result: 'ok' }))

  // Raise two tool calls, which the mod's tool.call hook counts
  await $.tool.call({ tool: 'Bash', command: 'ls' })
  await $.tool.call({ tool: 'Read', file_path: 'README.md' })

  // Run /tally and check the text its hook returns
  const answer = await $.command.run({ command: 'tally', args: '' })
  expect(answer.text).toBe('Claude has made 2 tool calls since this mod loaded')
})

Nella tua shell, esegui i test dalla directory first-mod:

claude plugin test

L'output nomina ogni test e se è passato, con tempi che variano da esecuzione a esecuzione:

tests/first-mod.test.ts:
(pass) /tally reports the tool calls the mod has seen [22.87ms]

 1 pass
 0 fail
Ran 1 test across 1 file. [0.19s]

Ogni $.tool.call è passato attraverso l'hook tool.call del mod, che ha aggiunto uno al suo conteggio e ha passato la chiamata allo stub. Nessun ls è stato eseguito e nessun file è stato letto. $.command.run è poi andato all'hook command.run del mod, e answer è l'oggetto che quell'hook ha restituito.

Il comando esce con stato 1 quando un test fallisce, quindi funziona in CI. Se i tuoi mod non riescono a caricarsi nella shell che lo esegue, stampa una riga che inizia con claude plugin test: hooks modules are turned off con il motivo, ed esce con stato 1.

Simula quello che Claude Code risponderebbe

Nessun modello, archivio o strumento viene eseguito in un test, quindi ovunque il tuo mod si aspetti una risposta da Claude Code, il test fornisce la risposta con uno stub. Una funzione di test riceve due argomenti per questo:

  • $: il $ del test, che sta dove Claude Code fa. Non è l'API dei mod che un hook riceve. Ognuno dei suoi metodi genera l'evento dello stesso nome, lo invia attraverso i tuoi hook del mod, e si risolve nel risultato: $.tool.call({ tool: 'Bash', command: 'ls' }) genera tool.call. $.command.run, $.prompt.submit, $.session.start, e $.turn.complete funzionano allo stesso modo, e $.classic.Stop e gli altri metodi $.classic generano un evento hook delle impostazioni. Un test non può generare direttamente una chiamata API dei mod come ui.close. Attivala attraverso il tuo mod, ad esempio premendo il pulsante che chiude il riquadro.
  • on: chiamalo per registrare gli stub, che sono hook che rispondono al posto di Claude Code. Nomina uno stub per una chiamata API dei mod senza il $., quindi uno stub registrato come store.get risponde al $.store.get del tuo mod. Quando il tuo mod chiama $.model.complete o $.store.get, uno stub fornisce la risposta.

Questo esempio simula una chiamata di modello. L'hook appartiene a un mod denominato grader, e gestisce un comando /grade che invia una frase a un modello e segnala se la risposta inizia con PASS. Il file contiene solo l'hook in prova, quindi il mod ha anche bisogno di un plugin.json e un hooks.json, come in Create a mod. Per digitare /grade in una sessione, il mod deve anche registrare il comando:

export function register(on) {
  on('command.run', { command: 'grade' }, async ($, e) => {
    // e.args is the text typed after /grade
    const reply = await $.model.complete({
      model: 'haiku',
      system: 'Grade the sentence. Start your reply with PASS or FAIL.',
      prompt: e.args,
    })
    const passed = reply.isAnswered && reply.text.startsWith('PASS')
    return { text: passed ? 'Passed' : 'Try again' }
  })
}

Questo test simula la chiamata del modello per controllare cosa fa l'hook con una risposta positiva:

import { expect, test } from 'claude-code/testing'

test('a passing grade is reported', async ($, on) => {
  // Answer the mod's $.model.complete call with a fixed reply, so no model runs
  on('model.complete', () => ({
    value: {
      isAnswered: true,
      text: 'PASS\nNice sentence.',
      usage: { input_tokens: 10, output_tokens: 5, cache_read_input_tokens: 0, cache_creation_input_tokens: 0 },
    },
  }))

  // Run /grade, which makes the mod call the model
  const answer = await $.command.run({ command: 'grade', args: 'The cat sat on the mat.' })
  expect(answer.text).toBe('Passed')
})

Il test passa perché il reply dell'hook è l'oggetto sotto value, il cui text inizia con PASS. Per controllare l'altro ramo, aggiungi un secondo test il cui stub restituisce un text che inizia con FAIL, e aspettati Try again.

Uno stub per una chiamata API dei mod restituisce un oggetto con un campo value, che contiene ciò a cui la chiamata si risolve nel tuo mod: { value: 7 } fa sì che $.store.get si risolva in 7. Uno stub per uno degli eventi di Claude Code, come turn.step o tool.call, restituisce il risultato proprio di quell'evento, come { result: 'ok' }. $.session.send e $.prompt.fill prendono anche il risultato dell'evento, come mostra la tabella. Guarda cosa restituisce uno stub mostra quale forma assume ogni nome comune. Due errori significano che uno stub è sbagliato o mancante. L'output di un test fallito include un blocco intitolato the engine reported:, e ogni errore appare lì:

  • returned neither { value } nor { deny }: uno stub per una chiamata API dei mod ha restituito un valore nudo
  • no implementation for seguito da un nome: il tuo mod ha fatto quella chiamata e nessuno stub la risponde

Il kit esporta anche mock in memoria che rispondono a un intero namespace per te. mock.clock(on) risponde a $.clock, mock.store(on, { count: 7 }) risponde a $.store da un archivio che inizia con quelle voci, e mock.env(on, { CI: 'true' }) risponde a $.env.get da quelle variabili. mock.clock restituisce un orologio mock che il tuo test avanza, quindi un test di un timer non aspetta. mock.store non restituisce nulla, quindi per controllare cosa il tuo mod ha salvato, scrivi i due stub store tu stesso come fa il test di disegno.

Segui le regole del test kit

Il test kit ha alcune regole proprie, e infrangerne una produce gli errori che i nuovi autori di test incontrano per primi:

  • Registra ogni stub prima della prima chiamata del test su $. Chiamare on dopo questo genera un errore come on("ui.render") after the test first called $.

  • session.start non viene eseguito da solo. Ogni test inizia con il tuo modulo appena caricato e nessuno dei suoi hook chiamato, quindi le variabili a livello di modulo mantengono i loro valori iniziali. Se un hook dipende da ciò che session.start configura, generalo per primo:

    // Answer the event after your hook passes it on with next(e)
    on('session.start', () => ({ cwd: '/work' }))
    // Answer the $.command.register call your hook makes
    on('command.register', () => ({ value: undefined }))
    // Raise the event, which runs your session.start hook
    await $.session.start({ surface: 'terminal', isInteractive: true, cwd: '/work' })
    

    Il secondo stub risponde alla chiamata $.command.register che un hook session.start come quello del tutorial fa. Senza di esso, quella chiamata rifiuta con no implementation for command.register e il kit salta il tuo hook, quindi nulla dopo la chiamata nell'hook viene eseguito. Il test non fallisce a quel punto. L'hook saltato è elencato sotto the engine reported: solo se un controllo successivo fallisce.

  • Un hook che restituisce next(e) ha bisogno di uno stub per rispondere. Quando il tuo hook ui.render restituisce next(e), ad esempio per non disegnare nulla mentre Claude è inattivo, montarlo fallisce con no implementation for ui.render. Registra uno stub che restituisce un elemento come dati semplici:

    // Stands for what Claude Code would draw at the site
    on('ui.render', () => ({ type: 'Text', props: {}, children: ['drawn by Claude Code'] }))
    

    Con lo stub registrato, il montaggio ha successo, e ui.find({ type: 'Text' }) restituisce quell'elemento ogni volta che il tuo hook ha restituito next(e).

  • Uno stub per turn.step è un generatore asincrono, e il test legge il flusso fino alla fine per ottenere il risultato:

    on('turn.step', async function* ($, e) {
      // Each yield is one piece of the model's streamed reply
      yield { kind: 'text', index: 0, text: 'ok' }
      // The return value is the result of the whole request
      return { turnId: e.turnId, index: e.index, answer: 'ok', toolUses: [], stopReason: 'end_turn', usage: null }
    })
    
    // Raise one request to the model, which runs your turn.step hook
    const stream = $.turn.step({ turnId: 't', index: 0, model: 'claude-test', messageCount: 1 })
    // Read every piece until the stream says it's done
    let step = await stream.next()
    while (step.done !== true) step = await stream.next()
    const result = step.value
    

    Quando il ciclo termina, result è l'oggetto che lo stub ha restituito, dopo che il tuo hook turn.step ha avuto la possibilità di cambiarlo. Qui result.answer è 'ok'.

  • Genera una chiamata di strumento con il nome dello strumento e gli argomenti come campi, come await $.tool.call({ tool: 'Bash', command: 'ls' }), e registra uno stub tool.call che restituisce { result }.

Guarda cosa restituisce uno stub

Ogni chiamata API dei mod che il tuo mod fa in un test ha bisogno di uno stub che risponda al posto di Claude Code, tranne i pochi che il kit risponde da solo: $.ui.invalidate e $.state chiama. Per le chiamate $.clock, usa mock.clock(on), altrimenti il $.clock.now() del tuo mod fallisce con no implementation for clock.now.

Questa tabella elenca quelli che i mod usano di più. La prima colonna è la chiamata che il tuo mod fa o l'evento che passa con next(e). La seconda è la funzione da passare a on con quel nome, quindi la riga $.store.get diventa on('store.get', ($, e) => ({ value: saved.get(e.key) })). Un '...' in uno stub contrassegna il testo che devi compilare:

Il tuo mod chiama o passa Stub
$.command.register, $.tool.register, $.ui.toast, $.ui.log, $.ui.status, $.ui.close, $.store.set () => ({ value: undefined }). Per ui.toast e ui.log, il testo è e.text.
$.store.get ($, e) => ({ value: saved.get(e.key) })
$.fs.read ($, e) => ({ value: e.path.endsWith('notes.md') ? '# Notes' : '' }). e.path arriva come percorso assoluto, quindi confronta con endsWith.
$.ui.open () => ({ value: { isPlaced: true } })
$.ui.ask Uno stub tool.call, perché la domanda la raggiunge come una chiamata allo strumento AskUserQuestion: ($, e) => ({ result: { answers: { [e.questions[0].question]: 'Run it' } } }). Controlla e.tool per primo se il tuo mod passa altre chiamate di strumento.
$.model.complete () => ({ value: { isAnswered: true, text: '...', usage } })
$.process.run ($, e) => ({ value: { exitCode: 0, stdout: '...', stderr: '' } }). e.argv è l'elenco degli argomenti e e.init contiene cwd e timeoutMs.
Qualsiasi chiamata API dei mod che dovrebbe fallire () => ({ deny: 'the reason' }), che fa sì che la chiamata rifiuti nel tuo mod. Uno stub che genera un'eccezione viene saltato invece.
session.start () => ({ cwd: '/work' })
turn.start ($, e) => ({ turnId: e.turnId })
tool.call () => ({ result: '...' })
turn.complete () => ({ text: '' }). Generalo con $.turn.complete({ turnId, answer, durationMs, isAborted: false, usage: null }).
prompt.submit ($, e) => ({ text: e.text })
prompt.fill () => ({ isFilled: true })
$.prompt.read () => ({ value: { text: '...', cursor: 0 } })
$.ui.copy () => ({ value: { isCopied: true } })
$.session.messages () => ({ value: [{ role: 'assistant', text: '...', toolUses: [] }] })
$.session.id, $.agent.list () => ({ value: 'abc123' }), () => ({ value: [] })
session.send () => ({ isDelivered: true }). e.to arriva come stringa anche quando il tuo mod ha passato { sessionId }.
session.receive ($, e) => ({ text: e.text }). Generalo con $.session.receive({ origin: { kind: 'peer-send-message' }, text }).
ui.render () => ({ type: 'Text', props: {}, children: ['...'] })

expect ha le asserzioni toBe, toEqual, toMatch, toMatchObject, toContain, toBeDefined, toBeUndefined, e toThrow, e .not prima di una qualsiasi di esse.

Testa un timer

Un mod che esegue lavoro su un timer ha bisogno di un orologio che il test controlla, così il test può spostare il tempo in avanti invece di aspettare. const clock = mock.clock(on) restituisce un orologio mock che inizia a 0 e si muove solo quando il tuo test lo muove. Per iniziare in un altro momento, passalo in millisecondi, come in mock.clock(on, { now: 5000 }). L'orologio ha questi metodi:

Metodo Cosa fa
await clock.advance(1000) Sposta il tempo in avanti di quel numero di millisecondi ed esegue ogni timer che scade
await clock.set(5000) Sposta il tempo in avanti a quel valore, come farebbe advance
clock.now() Restituisce il tempo, che è ciò a cui si risolve $.clock.now() del tuo mod
await clock.settle() Esegue i timer che sono già scaduti, come una catena di chiamate $.clock.after a ritardo zero, senza spostare il tempo
await clock.sleep(2000) All'interno di uno stub, fa sì che quello stub risponda solo una volta che il test ha avanzato così lontano, che è come simuli un modello o un processo lento

Questo hook appartiene a un mod denominato countdown, e gestisce un comando /countdown che accetta un numero di secondi, avvia un timer $.clock.every di un secondo, e mostra un toast a zero. Come con grader, il file contiene solo l'hook in prova e non registra il comando:

export function register(on) {
  on('command.run', { command: 'countdown' }, async ($, e) => {
    // e.args is the text typed after /countdown
    let left = Number(e.args)
    const timer = $.clock.every(1000, () => {
      left -= 1
      if (left === 0) {
        timer.cancel()
        $.ui.toast('Time is up')
      }
    })
    // Print nothing in the transcript
    return {}
  })
}

Questo test esegue /countdown 3 e sposta l'orologio mock, così controlla tre secondi di comportamento senza aspettare tre secondi:

import { expect, mock, test } from 'claude-code/testing'

test('the countdown ends with a toast', async ($, on) => {
  // Answer every $.clock call from a clock the test controls
  const clock = mock.clock(on)
  // Collect the text of each toast the mod shows
  const toasts: string[] = []
  on('ui.toast', ($, e) => {
    toasts.push(e.text)
    return { value: undefined }
  })

  await $.command.run({ command: 'countdown', args: '3' })
  // After two seconds the timer has fired twice, and no toast is due
  await clock.advance(2000)
  expect(toasts).toEqual([])
  // The third second brings the count to zero
  await clock.advance(1000)
  expect(toasts).toEqual(['Time is up'])
})

Il primo expect mostra che il toast non arriva presto, e il secondo mostra che arriva una volta. Ogni advance si risolve dopo che i timer che sono scaduti hanno eseguito, così il controllo sulla riga successiva vede il loro effetto.

Testa un disegno

Un test può disegnare uno dei siti di rendering del tuo mod, quindi premere, digitare in, e trovare gli elementi che ha disegnato. $.ui.mount disegna il sito attraverso l'hook ui.render del tuo mod e restituisce un handle con un metodo per ognuno di quelli. Per coprire diverse app in un test, imposta surface all'app per cui disegnare. Questo test apre il riquadro da Build a pane with tabs, cambia schede, preme il pulsante, e controlla il conteggio nel terminale e nell'app Desktop:

import { expect, test } from 'claude-code/testing'

// What Claude Code passes to a ui.render hook for this pane, apart from the app
const PANE = {
  plugin: 'hello-tabs',
  component: 'Pane',
  requestId: 'hello-tabs',
  viewport: { columns: 100, rows: 30 },
  props: {
    title: 'Hello tabs',
    isFocused: true,
    bodyColumns: 60,
    placement: 'inline',
    scroll: { offset: 0, bodyRows: 10 },
    view: {},
  },
} as const

test('the second tab counts presses and saves the count', async ($, on) => {
  // Stub $.store with a Map, so the test can read what the mod saved
  const saved = new Map<string, unknown>()
  on('store.get', ($, e) => ({ value: saved.get(e.key) }))
  on('store.set', ($, e) => {
    saved.set(e.key, e.value)
    return { value: undefined }
  })

  // Draw the pane once for each app
  for (const surface of ['terminal', 'desktop'] as const) {
    const ui = await $.ui.mount({ ...PANE, surface })
    // Press the buttons by the key the mod gave them
    await ui.press({ key: 'tab-two' })
    await ui.press({ key: 'more' })
    // The second tab's count line is in the drawing
    expect(await ui.find({ type: 'Text', text: /^Count: \d+$/ })).toBeDefined()
    await ui.unmount()
  }

  // One press in each app makes two
  expect(saved.get('count')).toBe(2)
})

Nella tua shell, esegui claude plugin test dalla directory hello-tabs. Il test passa quando entrambe le app disegnano la riga di conteggio e il mod ha salvato 2. Il conteggio si trasporta dalla prima app alla seconda perché entrambi i montaggi usano lo stesso modulo caricato.

L'handle che $.ui.mount restituisce ha questi metodi, che indirizzano gli elementi per la key che hai dato loro:

Metodo Cosa fa
press({ key: 'more' }) Preme il Button con quella key
input({ key: 'new-note', text: 'buy milk' }) Digita il testo nell'Input con quella key e preme Invio. Aggiungi kind: 'change' per digitare senza inviare.
select({ key: 'size', value: 'large' }) Sceglie l'opzione con quel valore nel Select con quella key
find({ key: 'more' }) o find({ type: 'Text', text: 'Count: 2' }) Restituisce il primo elemento corrispondente come { type, props, children }, o undefined. text può essere una stringa o un'espressione regolare.
unmount() Rimuove il disegno

Ogni metodo si risolve dopo che il tuo handler ha finito, così puoi controllare il risultato sulla riga successiva. Imposta props a ciò che Claude Code passerebbe per quel sito. La tabella dei siti di rendering elenca i prop di ogni sito, e i tipi per la tua build hanno i loro tipi.

Un test di disegno controlla l'albero che il tuo hook restituisce e se è valido per quell'app. Non controlla come l'app lo dipinge, quindi guarda un nuovo layout in una sessione reale anche.

Testa un disegno dopo `/clear`

Ogni test inizia con ogni valore $.state al suo default, che è come /clear li lascia. Per testare cosa fa il tuo mod dopo, salta session.start, genera classic.SessionStart con source: 'clear', e controlla cosa disegna il tuo mod.

Questo test controlla il modulo da Load a saved value again after /clear. Aggiungilo al file da Test a drawing, dove PANE è definito. Il primo test di quel file si aspetta che il pulsante salvi il conteggio, come il pulsante in Save from more than one session fa:

test('the saved count comes back after /clear', async ($, on) => {
  // The store already holds a count of 7
  on('store.get', () => ({ value: 7 }))
  // Answer the event after your hook passes it on with next(e)
  on('classic.SessionStart', () => ({}))

  // Raise the event that fires after /clear, which runs your hook
  await $.classic.SessionStart({ source: 'clear' })

  const ui = await $.ui.mount({ ...PANE, surface: 'terminal' })
  await ui.press({ key: 'tab-two' })
  // The pane shows the stored count, not the default of 0
  expect(await ui.find({ type: 'Text', text: 'Count: 7' })).toBeDefined()
})

Il test passa quando il tuo hook classic.SessionStart ha copiato il 7 salvato in $.state prima che il riquadro disegni. Senza quell'hook nel tuo modulo, il riquadro disegna Count: 0, find restituisce undefined, e il test fallisce a toBeDefined.

Testa un mod che giudica altri mod

Un mod che la tua organizzazione elenca in prependPlugins può rifiutare un altro mod prima che carichi. Per testarne uno, imposta il tier del tuo mod e dai al test un secondo mod per il tuo da ammettere o rifiutare:

  • tier: chiamalo una volta in cima al file di test, come in tier('prepend'), per caricare il tuo mod come prepend, append, o builtin, il suo posto nell'ordine in cui i mod vengono eseguiti. Senza di esso, il tuo mod carica come user.
  • plugins: passa a test un oggetto opzioni davanti al corpo del test. Il suo array plugins contiene mod che scrivi inline, ognuno con un name e una funzione register. Per caricare uno da qualche parte diversa da user, aggiungi tier ad esso.

Questo file di test carica il mod di policy dalla pagina admin per primo. Controlla che il mod di policy rifiuti un mod che avvia un processo e ammetta uno che non lo fa:

import { expect, test, tier } from 'claude-code/testing'

// Load the mod under test ahead of every other mod
tier('prepend')

// A second mod whose code calls $.process.run, which the policy blocks
const runner = {
  name: 'runner',
  register(on) {
    on('tool.call', async ($, e, next) => {
      await $.process.run(['ls'])
      return { result: 'runner answered' }
    })
  },
}

// A second mod that calls nothing the policy blocks
const reader = {
  name: 'reader',
  register(on) {
    on('tool.call', async ($, e, next) => {
      return { result: 'reader answered' }
    })
  },
}

test('refuses a mod that starts a process', { plugins: [runner] }, async ($, on) => {
  on('tool.call', () => ({ result: 'claude code answered' }))
  let message = ''
  try {
    // The first call on $ loads the mods, so the refusal is thrown here
    await $.tool.call({ tool: 'Bash', command: 'ls' })
  } catch (error) {
    message = error.message
  }
  expect(message).toBe('runner: refused by acme-guard: Acme policy: mods may not call process.run')
})

test('admits a mod that starts no process', { plugins: [reader] }, async ($, on) => {
  on('tool.call', () => ({ result: 'claude code answered' }))
  const out = await $.tool.call({ tool: 'Bash', command: 'ls' })
  // The answer comes from reader, which shows that it loaded
  expect(out).toEqual({ result: 'reader answered' })
})

Nella tua shell, esegui claude plugin test dalla directory acme-guard. Entrambi i test passano con il mod di policy come mostra la pagina admin.

Il kit carica ogni mod alla prima chiamata del test su $. Quando il tuo mod ne rifiuta uno, quella chiamata genera un'eccezione, e il messaggio nomina il mod rifiutato, il mod che lo ha rifiutato, e il tuo motivo. Nel secondo test nulla viene rifiutato, quindi reader risponde alla chiamata di strumento prima che raggiunga lo stub.

Passaggi successivi