Testare un mod
Scrivi test automatizzati per un mod di Claude Code che attivano eventi, simulano le risposte di Claude Code e premono pulsanti, senza sessione, accesso o rete.
Puoi scrivere test automatizzati per un mod ed eseguirli dalla tua shell con claude plugin test. Un test attiva gli eventi gestiti dai tuoi hook e verifica cosa hanno fatto gli hook, così individui un problema prima che raggiunga una sessione. Il primo esempio testa il mod di Creare un mod.
Scrivere un test
Un test carica il tuo mod, invia eventi attraverso i suoi hook come farebbe Claude Code e verifica cosa hanno fatto gli hook, senza una sessione, un accesso o una rete. Esegui i test dalla tua shell con claude plugin test, e ogni file di test importa il test kit, una libreria di test nel modulo claude-code/testing.
Dai a ogni file di test un nome che termini in .test.ts, come first-mod.test.ts, e salvalo in qualsiasi punto della directory del plugin. Ogni file di test ha bisogno di almeno un test(), altrimenti l'esecuzione fallisce con declares no test(): nothing ran. Un file di test può importare i file del tuo mod e gli helper .ts adiacenti, quindi puoi scrivere unit test per funzioni semplici, come le regole di un gioco, senza il kit.
Questo test attiva due chiamate agli strumenti, esegue il comando /tally da Creare un mod e verifica che la risposta le conti entrambe. La sua prima riga è uno stub, che risponde alle chiamate agli strumenti al posto di Claude Code. Salvalo come first-mod/tests/first-mod.test.ts:
import { expect, test } from 'claude-code/testing'
test('/tally reports the tool calls the mod has seen', async ($, on) => {
// Answer each tool call in Claude Code's place, so no tool runs
on('tool.call', () => ({ result: 'ok' }))
// Fire two tool calls, which the mod's tool.call hook counts
await $.tool.call({ tool: 'Bash', command: 'ls' })
await $.tool.call({ tool: 'Read', file_path: 'README.md' })
// Run /tally and check the text its hook returns
const answer = await $.command.run({ command: 'tally', args: '' })
expect(answer.text).toBe('Claude has made 2 tool calls since this mod loaded')
})
Nella tua shell, esegui i test dalla directory first-mod:
claude plugin test
L'output indica il nome di ogni test e se è stato superato, con tempi che variano da un'esecuzione all'altra:
tests/first-mod.test.ts:
(pass) /tally reports the tool calls the mod has seen [22.87ms]
1 pass
0 fail
Ran 1 test across 1 file. [0.19s]
Ogni $.tool.call è passata attraverso l'hook tool.call del mod, che ha aggiunto uno al suo conteggio e ha passato la chiamata allo stub. Non è stato eseguito alcun ls e non è stato letto alcun file. $.command.run è poi passato all'hook command.run del mod, e answer è l'oggetto restituito da quell'hook.
Il comando termina con stato 1 quando un test fallisce, quindi funziona in CI. Se i tuoi mod non possono essere caricati nella shell che lo esegue, stampa una riga che inizia con claude plugin test: hooks modules are turned off con il motivo, e termina con stato 1.
Creare stub per ciò che Claude Code risponderebbe
In un test non vengono eseguiti modelli, store o strumenti, quindi ovunque il tuo mod si aspetti una risposta da Claude Code, il test fornisce la risposta con uno stub. Una funzione di test riceve due argomenti a questo scopo:
$: il$proprio del test, che agisce come Claude Code. Non è la mods API che riceve un hook. Ciascuno dei suoi metodi attiva l'evento con lo stesso nome, lo invia attraverso gli hook del tuo mod e si risolve nel risultato:$.tool.call({ tool: 'Bash', command: 'ls' })attivatool.call.$.command.run,$.prompt.submit,$.session.starte$.turn.completefunzionano allo stesso modo, e$.classic.Stope gli altri metodi$.classicattivano un evento degli hook delle impostazioni. Un test non può attivare direttamente una chiamata della mods API comeui.close. Attivala tramite il tuo mod, ad esempio premendo il pulsante che chiude il riquadro.on: chiamalo per registrare gli stub, che sono hook che rispondono al posto di Claude Code. Assegna a uno stub per una chiamata della mods API il nome senza$., così uno stub registrato comestore.getrisponde al$.store.getdel tuo mod. Quando il tuo mod chiama$.model.completeo$.store.get, uno stub fornisce la risposta.
Questo esempio crea uno stub per una chiamata a un modello. L'hook appartiene a un mod chiamato grader e gestisce un comando /grade che invia una frase a un modello e riporta se la risposta inizia con PASS. Il file contiene solo l'hook sotto test, quindi il mod ha bisogno anche di un plugin.json e di un hooks.json, come in Creare un mod. Per digitare /grade in una sessione, il mod deve anche registrare il comando:
export function register(on) {
on('command.run', { command: 'grade' }, async ($, e) => {
// e.args is the text typed after /grade
const reply = await $.model.complete({
model: 'haiku',
system: 'Grade the sentence. Start your reply with PASS or FAIL.',
prompt: e.args,
})
const passed = reply.isAnswered && reply.text.startsWith('PASS')
return { text: passed ? 'Passed' : 'Try again' }
})
}
Questo test crea uno stub per la chiamata al modello per verificare cosa fa l'hook con una risposta positiva:
import { expect, test } from 'claude-code/testing'
test('a passing grade is reported', async ($, on) => {
// Answer the mod's $.model.complete call with a fixed reply, so no model runs
on('model.complete', () => ({
value: {
isAnswered: true,
text: 'PASS\nNice sentence.',
usage: { input_tokens: 10, output_tokens: 5, cache_read_input_tokens: 0, cache_creation_input_tokens: 0 },
},
}))
// Run /grade, which makes the mod call the model
const answer = await $.command.run({ command: 'grade', args: 'The cat sat on the mat.' })
expect(answer.text).toBe('Passed')
})
Il test viene superato perché il reply dell'hook è l'oggetto sotto value, il cui text inizia con PASS. Per verificare l'altro ramo, aggiungi un secondo test il cui stub restituisce un text che inizia con FAIL, e aspettati Try again.
Uno stub per una chiamata della mods API restituisce un oggetto con un campo value, che contiene ciò in cui la chiamata si risolve nel tuo mod: { value: 7 } fa sì che $.store.get si risolva in 7. Uno stub per uno degli eventi di Claude Code, come turn.step o tool.call, restituisce il risultato proprio di quell'evento, come { result: 'ok' }. Anche $.session.send e $.prompt.fill accettano il risultato del loro evento, come mostra la tabella. Consultare cosa restituisce uno stub mostra quale forma assume ciascun nome comune. Questi errori indicano che uno stub è errato o mancante. L'output di un test fallito include un blocco intitolato the engine reported:, e ogni errore compare lì:
returned neither { value } nor { deny }: uno stub per una chiamata della mods API ha restituito un valore sempliceno implementation forseguito da un nome: il tuo mod ha effettuato quella chiamata e nessuno stub risponde
Il kit esporta anche mock in memoria che rispondono per te a un intero namespace. mock.clock(on) risponde a $.clock, mock.store(on, { count: 7 }) risponde a $.store da uno store che inizia con quelle voci, e mock.env(on, { CI: 'true' }) risponde a $.env.get da quelle variabili. mock.clock restituisce un orologio mock che il tuo test fa avanzare, così un test di un timer non deve attendere. mock.store non restituisce nulla, quindi per verificare cosa ha salvato il tuo mod, scrivi tu stesso i due stub store come fa il test di disegno.
Seguire le regole del test kit
Il test kit ha alcune regole proprie, e violarne una produce gli errori in cui si imbattono per primi i nuovi autori di test:
-
Registra ogni stub prima della prima chiamata del test su
$. Chiamareondopo di essa genera un errore comeon("ui.render") after the test first called $. -
session.startnon viene eseguito da solo. Ogni test inizia con il tuo modulo appena caricato e nessuno dei suoi hook chiamato, quindi le variabili a livello di modulo mantengono i loro valori iniziali. Se un hook dipende da ciò chesession.startimposta, attivalo per primo:// Answer the event after your hook passes it on with next(e) on('session.start', () => ({ cwd: '/work' })) // Answer the $.command.register call your hook makes on('command.register', () => ({ value: undefined })) // Fire the event, which runs your session.start hook await $.session.start({ surface: 'terminal', isInteractive: true, cwd: '/work' })Il secondo stub risponde alla chiamata
$.command.registereffettuata da un hooksession.startcome quello del tutorial. Senza di esso, quella chiamata viene rifiutata conno implementation for command.registere il kit salta il tuo hook, quindi nulla di ciò che segue la chiamata nell'hook viene eseguito. Il test non fallisce in quel punto. L'hook saltato viene elencato sottothe engine reported:solo se una verifica successiva fallisce. -
Un hook che restituisce
next(e)ha bisogno di uno stub che risponda. Quando il tuo hookui.renderrestituiscenext(e), ad esempio per non disegnare nulla mentre Claude è inattivo, montarlo fallisce conno implementation for ui.render. Registra uno stub che restituisce un elemento come dati semplici:// Stands for what Claude Code would draw at the site on('ui.render', () => ({ type: 'Text', props: {}, children: ['drawn by Claude Code'] }))Con lo stub registrato, il montaggio riesce, e
ui.find({ type: 'Text' })restituisce quell'elemento ogni volta che il tuo hook ha restituitonext(e). -
Uno stub per
turn.stepè un generatore asincrono, e il test legge lo stream fino alla fine per ottenere il risultato:on('turn.step', async function* ($, e) { // Each yield is one piece of the model's streamed reply yield { kind: 'text', index: 0, text: 'ok' } // The return value is the result of the whole request return { turnId: e.turnId, index: e.index, answer: 'ok', toolUses: [], stopReason: 'end_turn', usage: null } }) // Fire one request to the model, which runs your turn.step hook const stream = $.turn.step({ turnId: 't', index: 0, model: 'claude-test', messageCount: 1 }) // Read every piece until the stream says it's done let step = await stream.next() while (step.done !== true) step = await stream.next() const result = step.valueQuando il ciclo termina,
resultè l'oggetto restituito dallo stub, dopo che il tuo hookturn.stepha avuto la possibilità di modificarlo. Quiresult.answerè'ok'. -
Attiva una chiamata a uno strumento con il nome e gli argomenti dello strumento come campi, come
await $.tool.call({ tool: 'Bash', command: 'ls' }), e registra uno stubtool.callche restituisce{ result }.
Consultare cosa restituisce uno stub
Ogni chiamata della mods API che il tuo mod effettua in un test ha bisogno di uno stub che risponda al posto di Claude Code, tranne le poche a cui risponde il kit stesso: le chiamate $.ui.invalidate e $.state. Per le chiamate $.clock, usa mock.clock(on), altrimenti il $.clock.now() del tuo mod fallisce con no implementation for clock.now.
Questa tabella elenca quelle che i mod usano più spesso. La prima colonna è la chiamata che il tuo mod effettua o l'evento che passa con next(e). La seconda è la funzione da passare a on con quel nome, così la riga $.store.get diventa on('store.get', ($, e) => ({ value: saved.get(e.key) })). Un '...' in uno stub indica un testo da compilare:
| Il tuo mod chiama o passa | Stub |
|---|---|
$.command.register, $.tool.register, $.ui.toast, $.ui.log, $.ui.status, $.ui.close, $.store.set |
() => ({ value: undefined }). Per ui.toast e ui.log, il testo è e.text. |
$.store.get |
($, e) => ({ value: saved.get(e.key) }) |
$.fs.read |
($, e) => ({ value: e.path.endsWith('notes.md') ? '# Notes' : '' }). e.path arriva come percorso assoluto, quindi confronta con endsWith. |
$.ui.open |
() => ({ value: { isPlaced: true } }) |
$.ui.ask |
Uno stub tool.call, perché la domanda gli arriva come chiamata allo strumento AskUserQuestion: ($, e) => ({ result: { answers: { [e.questions[0].question]: 'Run it' } } }). Verifica prima e.tool se il tuo mod passa altre chiamate agli strumenti. |
$.model.complete |
() => ({ value: { isAnswered: true, text: '...', usage } }) |
$.process.run |
($, e) => ({ value: { exitCode: 0, stdout: '...', stderr: '' } }). e.argv è l'elenco degli argomenti e e.init contiene cwd e timeoutMs. |
| Qualsiasi chiamata della mods API che dovrebbe fallire | () => ({ deny: 'the reason' }), che fa sì che la chiamata venga rifiutata nel tuo mod. Uno stub che genera un'eccezione viene invece saltato. |
session.start |
() => ({ cwd: '/work' }) |
turn.start |
($, e) => ({ turnId: e.turnId }) |
tool.call |
() => ({ result: '...' }) |
turn.complete |
() => ({ text: '' }). Attivalo con $.turn.complete({ turnId, answer, durationMs, isAborted: false, usage: null }). |
prompt.submit |
($, e) => ({ text: e.text }) |
prompt.fill |
() => ({ isFilled: true }) |
$.prompt.read |
() => ({ value: { text: '...', cursor: 0 } }) |
$.ui.copy |
() => ({ value: { isCopied: true } }) |
$.session.messages |
() => ({ value: [{ role: 'assistant', text: '...', toolUses: [] }] }) |
$.session.id, $.agent.list |
() => ({ value: 'abc123' }), () => ({ value: [] }) |
session.send |
() => ({ isDelivered: true }). e.to arriva come stringa anche quando il tuo mod ha passato { sessionId }. |
session.receive |
($, e) => ({ text: e.text }). Attivalo con $.session.receive({ origin: { kind: 'peer-send-message' }, text }). |
ui.render |
() => ({ type: 'Text', props: {}, children: ['...'] }) |
expect dispone delle asserzioni toBe, toEqual, toMatch, toMatchObject, toContain, toBeDefined, toBeUndefined e toThrow, e di .not prima di ciascuna di esse.
Testare un timer
Un mod che esegue lavoro su un timer ha bisogno di un orologio controllato dal test, così il test può far avanzare il tempo invece di aspettare. const clock = mock.clock(on) restituisce un orologio mock che parte da 0 e avanza solo quando è il tuo test a farlo avanzare. Per partire da un altro istante, passalo in millisecondi, come in mock.clock(on, { now: 5000 }). L'orologio ha questi metodi:
| Metodo | Cosa fa |
|---|---|
await clock.advance(1000) |
Fa avanzare il tempo di quel numero di millisecondi ed esegue ogni timer che arriva a scadenza |
await clock.set(5000) |
Fa avanzare il tempo fino a quel valore, come farebbe advance |
clock.now() |
Restituisce il tempo, che è il valore a cui si risolve $.clock.now() del tuo mod |
await clock.settle() |
Esegue i timer già scaduti, come una catena di chiamate $.clock.after con ritardo zero, senza far avanzare il tempo |
await clock.sleep(2000) |
All'interno di uno stub, fa sì che quello stub risponda solo quando il test è avanzato fino a quel punto, ed è così che simuli un modello o un processo lento |
Questo hook appartiene a un mod chiamato countdown e gestisce un comando /countdown che accetta un numero di secondi, avvia un timer $.clock.every di un secondo e mostra un toast quando arriva a zero. Come per grader, il file contiene solo l'hook sotto test e non registra il comando:
export function register(on) {
on('command.run', { command: 'countdown' }, async ($, e) => {
// e.args is the text typed after /countdown
let left = Number(e.args)
const timer = $.clock.every(1000, () => {
left -= 1
if (left === 0) {
timer.cancel()
$.ui.toast('Time is up')
}
})
// Print nothing in the transcript
return {}
})
}
Questo test esegue /countdown 3 e fa avanzare l'orologio mock, così verifica tre secondi di comportamento senza aspettare tre secondi:
import { expect, mock, test } from 'claude-code/testing'
test('the countdown ends with a toast', async ($, on) => {
// Answer every $.clock call from a clock the test controls
const clock = mock.clock(on)
// Collect the text of each toast the mod shows
const toasts: string[] = []
on('ui.toast', ($, e) => {
toasts.push(e.text)
return { value: undefined }
})
await $.command.run({ command: 'countdown', args: '3' })
// After two seconds the timer has fired twice, and no toast is due
await clock.advance(2000)
expect(toasts).toEqual([])
// The third second brings the count to zero
await clock.advance(1000)
expect(toasts).toEqual(['Time is up'])
})
Il primo expect mostra che il toast non arriva in anticipo, e il secondo mostra che arriva una sola volta. Ogni advance si risolve dopo che i timer arrivati a scadenza sono stati eseguiti, quindi il controllo alla riga successiva ne vede l'effetto.
Testare un disegno
Un test può disegnare uno dei punti di rendering del tuo mod, quindi premere, digitare e trovare gli elementi che ha disegnato. $.ui.mount disegna il punto tramite l'hook ui.render del tuo mod e restituisce un handle con un metodo per ciascuna di queste azioni. Per coprire più app in un solo test, imposta surface sull'app per cui disegnare. Questo test apre il pannello di Creare un pannello con schede, cambia scheda, preme il pulsante e controlla il conteggio nel terminale e nell'app Desktop:
import { expect, test } from 'claude-code/testing'
// What Claude Code passes to a ui.render hook for this pane, apart from the app
const PANE = {
plugin: 'hello-tabs',
component: 'Pane',
requestId: 'hello-tabs',
viewport: { columns: 100, rows: 30 },
props: {
title: 'Hello tabs',
isFocused: true,
bodyColumns: 60,
placement: 'inline',
scroll: { offset: 0, bodyRows: 10 },
view: {},
},
} as const
test('the second tab counts presses and saves the count', async ($, on) => {
// Stub $.store with a Map, so the test can read what the mod saved
const saved = new Map<string, unknown>()
on('store.get', ($, e) => ({ value: saved.get(e.key) }))
on('store.set', ($, e) => {
saved.set(e.key, e.value)
return { value: undefined }
})
// Draw the pane once for each app
for (const surface of ['terminal', 'desktop'] as const) {
const ui = await $.ui.mount({ ...PANE, surface })
// Press the buttons by the key the mod gave them
await ui.press({ key: 'tab-two' })
await ui.press({ key: 'more' })
// The second tab's count line is in the drawing
expect(await ui.find({ type: 'Text', text: /^Count: \d+$/ })).toBeDefined()
await ui.unmount()
}
// One press in each app makes two
expect(saved.get('count')).toBe(2)
})
Nella tua shell, esegui claude plugin test dalla directory hello-tabs. Il test ha esito positivo quando entrambe le app disegnano la riga del conteggio e il mod ha salvato 2. Il conteggio passa dalla prima app alla seconda perché entrambi i mount usano lo stesso modulo caricato.
L'handle restituito da $.ui.mount ha questi metodi, che individuano gli elementi tramite la key che hai assegnato loro:
| Metodo | Cosa fa |
|---|---|
press({ key: 'more' }) |
Preme il Button con quella key |
input({ key: 'new-note', text: 'buy milk' }) |
Digita il testo nell'Input con quella key e preme Invio. Aggiungi kind: 'change' per digitare senza inviare. |
select({ key: 'size', value: 'large' }) |
Sceglie l'opzione con quel valore nel Select con quella key |
find({ key: 'more' }) o find({ type: 'Text', text: 'Count: 2' }) |
Restituisce il primo elemento corrispondente come { type, props, children }, oppure undefined. text può essere una stringa o un'espressione regolare. |
unmount() |
Rimuove il disegno |
Ogni metodo si risolve dopo che il tuo handler ha terminato, quindi puoi controllare il risultato nella riga successiva. Imposta props su ciò che Claude Code passerebbe per quel punto. La tabella dei punti di rendering elenca le prop di ciascun punto, e i tipi per la tua build ne contengono i tipi.
Un test di disegno controlla l'albero restituito dal tuo hook e se è valido per quell'app. Non controlla come l'app lo visualizza, quindi verifica un nuovo layout anche in una sessione reale.
Testare un disegno dopo `/clear`
Ogni test inizia con ogni valore di $.state impostato al suo default, che è lo stato in cui li lascia /clear. Per testare cosa fa il tuo mod dopo, salta session.start, attiva classic.SessionStart con source: 'clear' e controlla cosa disegna il tuo mod.
Questo test controlla il modulo di Caricare di nuovo un valore salvato dopo /clear. Aggiungilo al file di Testare un disegno, dove è definito PANE. Il primo test di quel file si aspetta che il pulsante salvi il conteggio, come fa il pulsante in Salvare da più di una sessione:
test('the saved count comes back after /clear', async ($, on) => {
// The store already holds a count of 7
on('store.get', () => ({ value: 7 }))
// Answer the event after your hook passes it on with next(e)
on('classic.SessionStart', () => ({}))
// Fire the event that follows /clear, which runs your hook
await $.classic.SessionStart({ source: 'clear' })
const ui = await $.ui.mount({ ...PANE, surface: 'terminal' })
await ui.press({ key: 'tab-two' })
// The pane shows the stored count, not the default of 0
expect(await ui.find({ type: 'Text', text: 'Count: 7' })).toBeDefined()
})
Il test ha esito positivo quando il tuo hook classic.SessionStart ha copiato il 7 memorizzato in $.state prima che il pannello venga disegnato. Senza quell'hook nel tuo modulo, il pannello disegna Count: 0, find restituisce undefined e il test fallisce su toBeDefined.
Testare un mod di criteri
Un mod che la tua organizzazione elenca in prependPlugins può rifiutare un altro mod prima che venga caricato. Per testarne uno, imposta il livello del tuo mod e fornisci al test un secondo mod che il tuo possa consentire o rifiutare:
tier: chiamalo una volta all'inizio del file di test, come intier('prepend'), per caricare il tuo mod comeprepend,appendobuiltin, ovvero la sua posizione nell'ordine in cui vengono eseguiti i mod. Senza di esso, il tuo mod viene caricato comeuser.plugins: passa atestun oggetto di opzioni prima del corpo del test. Il suo arraypluginscontiene i mod che scrivi inline, ciascuno con unnamee una funzioneregister. Per caricarne uno in una posizione diversa dauser, aggiungitieral mod.
Questo file di test carica per primo il mod di criteri della pagina di amministrazione. Verifica che il mod di criteri rifiuti un mod che avvia un processo e consenta un mod che non lo fa:
import { expect, test, tier } from 'claude-code/testing'
// Load the mod under test ahead of every other mod
tier('prepend')
// A second mod whose code calls $.process.run, which the policy blocks
const runner = {
name: 'runner',
register(on) {
on('tool.call', async ($, e, next) => {
await $.process.run(['ls'])
return { result: 'runner answered' }
})
},
}
// A second mod that calls nothing the policy blocks
const reader = {
name: 'reader',
register(on) {
on('tool.call', async ($, e, next) => {
return { result: 'reader answered' }
})
},
}
test('refuses a mod that starts a process', { plugins: [runner] }, async ($, on) => {
on('tool.call', () => ({ result: 'claude code answered' }))
let message = ''
try {
// The first call on $ loads the mods, so the refusal is thrown here
await $.tool.call({ tool: 'Bash', command: 'ls' })
} catch (error) {
message = error.message
}
expect(message).toBe('runner: refused by acme-guard: Acme policy: mods may not call process.run')
})
test('admits a mod that starts no process', { plugins: [reader] }, async ($, on) => {
on('tool.call', () => ({ result: 'claude code answered' }))
const out = await $.tool.call({ tool: 'Bash', command: 'ls' })
// The answer comes from reader, which shows that it loaded
expect(out).toEqual({ result: 'reader answered' })
})
Nella tua shell, esegui claude plugin test dalla directory acme-guard. Entrambi i test vengono superati con il mod di criteri così come appare nella pagina di amministrazione.
Il kit carica ogni mod alla prima chiamata del test su $. Quando il tuo mod ne rifiuta uno, quella chiamata genera un errore e il messaggio indica il mod rifiutato, il mod che lo ha rifiutato e il tuo motivo. Nel secondo test non viene rifiutato nulla, quindi reader risponde alla chiamata allo strumento prima che raggiunga lo stub.
Passaggi successivi
- Risolvere i problemi di un mod: scopri perché un mod non fa nulla in una sessione
- Riferimento dei mod: input e risultato di ogni evento, per scrivere stub