The uiv.* API: JS Script Macros
Modern JavaScript Macros — API Overview
A Ui.Vision macro can be a modern JavaScript program: instead of a table of commands, you write real code against the uiv.* API. Everything JavaScript offers — loops, conditions, functions, try/catch, template literals — combines with everything Ui.Vision offers: DOM automation, computer vision, OCR, AI, desktop input. This page is the API overview. Coming from classic table macros? See the Selenium-derived commands vs uiv.* conversion table.
uiv.goto('https://en.wikipedia.org');
uiv.page.fill('id=searchInput', 'Solar cell');
uiv.browser.type('${KEY_ENTER}', {nav: true});
const h1 = uiv.$('css=#firstHeading'); // auto-waits, throws if it never appears
uiv.log(`Landed on: ${h1.text}`, 'green');
Two ideas carry the whole API: finders answer "where is it?" and return a match
{x, y, rect, text, value, ...}; input tiers act on a match (or locator) and differ
in how the input reaches the page. Every finder auto-waits up to !TIMEOUT_WAIT and then throws — finding an element IS the wait.
Finders — Where Is It?
| Method | What it does | Example |
|---|---|---|
uiv.$(locator) |
First DOM match. Finds in ALL frames (even cross-origin) and open shadow roots — no selectFrame needed. Locators: css= id= name= link= xpath= | uiv.$('css=#buy') |
uiv.$$(locator) |
All DOM matches, as an array | uiv.$$('css=tr').length |
uiv.getByRole(role, {name})uiv.getByLabel(text)uiv.getByText(text)uiv.getByPlaceholder / getByTitle / getByAltText / getByTestId |
Finders by accessible name, with Playwright's names and semantics: what the AI's page tree (and a screen reader) calls the element is the locator — no css/xpath to invent, and it survives re-renders. Default match is a case-insensitive substring; {exact: true} is the whole string; {all: true} returns every match. A match is a locator for every action tier. |
uiv.page.click(uiv.getByRole('button', {name: 'Rechnung erstellen'}))uiv.page.fill(uiv.getByLabel('PLZ'), '69190') |
match.innerText() .inputValue() .isVisible() .isChecked() .isEnabled() .getAttribute(name) |
Playwright's questions, answered by a DOM match from find time (like the universal fields .text, .value, .rect, .x/.y that every match has). An OCR or image match — browser or desktop — has no DOM and says so when asked. |
if (uiv.getByLabel('AGB').isChecked()) ... |
uiv.findImage(png, opts) |
First computer-vision match of an image on the visible page. Options: {minScore, area, scope} | uiv.findImage('buy.png', {minScore: 0.8}) |
uiv.ocr.findText(text) |
First OCR match of rendered text — works on canvas, PDFs, images. ? and * wildcards per word | uiv.ocr.findText('Check*ut') |
uiv.ai.find(question) |
The configured AI locates the element from a screenshot and returns a match | uiv.ai.find('the blue Buy button') |
uiv.offset(match, dx, dy) |
A new match at a fixed offset from another — the composed form of the classic relative targets | uiv.offset(uiv.ocr.findText('Email:'), 120, 0) |
uiv.findElements / uiv.findImages / uiv.ocr.findTexts |
Long forms returning every match, with options like {required: false, timeout: 2} for optional elements (null/[] instead of throwing) | uiv.findImages('row.png', {area: uiv.$('css=#list')}) |
Naming rule: a feature that Playwright also has carries Playwright's name in uiv — uiv.goto, uiv.evaluate, uiv.page.fill, uiv.page.selectOption, uiv.browser.hover, uiv.browser.press('Control+S') and the getBy* finders above. Features that also exist on the desktop (uiv.screenshot, .text, .rect, the desktop tier itself) keep their scope-neutral names, because Playwright is browser-only. The older spellings uiv.open, uiv.eval, uiv.page.type, uiv.page.select and uiv.browser.move still run.
Acting — Three Input Tiers
Every input call names its tier, because how the input reaches the page decides whether it works.
A visual click is always explicit: uiv.browser.click(uiv.findImage('buy.png')).
| uiv.page.* | uiv.browser.* | uiv.desktop.* | |
|---|---|---|---|
| Methods | click, fill(locator, text), selectOption, scrollIntoViewIfNeeded |
click, type(text), press, mouse.move/wheel, keyboard.down/up |
mouse.click/move/down/up/wheel; keyboard.type/press/insertText/down/up |
| Event type | Synthetic DOM events (like Selenium) | Trusted browser input via the debugger API (CDP) | Real OS input (XModule) |
| XModule needed | No | No | Yes |
| Coordinates | Viewport CSS pixels | Viewport CSS pixels | Screen pixels ({scope: 'desktop'} finders) |
| Browser in background OK | Yes | Yes | No — real mouse moves |
| Best for | Form filling, fast bulk work | Canvas apps, drag & drop, sites that ignore synthetic clicks | OS dialogs, other applications, desktop automation |
| Browsers | All | Chrome, Edge (not Firefox) | All (XModule: Windows, Mac, Linux) |
Dragging is press, move, release: uiv.browser.down(start) holds the button,
every uiv.browser.hover while it is held drags, uiv.browser.up(end) releases —
sliders, drag handles and canvas drawing all work this way.
Keystrokes go to the focused element and support the ${KEY_...} codes:
uiv.browser.type('${KEY_ENTER}', {nav: true}) submits and waits for the navigation it causes.
Text without keystrokes: uiv.desktop.keyboard.insertText(text, {chord, settle, focus})
(Playwright's keyboard.insertText) puts the text on the OS clipboard and presses one paste
chord (Control+V; Meta+V on macOS; pass chord: 'Control+Shift+V' for a Linux terminal), so no keyboard layout, dead key or
AltGr is involved. Use it for text that keystrokes mangle: a window on the far side of a remote-control viewer (AnyDesk, TeamViewer relay
keystrokes through the far side's layout and drop characters like | that the local layout cannot encode for it), a foreign-layout
guest, a field that rejects synthetic keys. settle (ms, default 1000) gives the clipboard sync time; focus is a point or
finder match clicked after the write and before the paste (a VM pulls the host clipboard when its window is clicked, so the order matters).
Key names are not interpreted in pasted text; press Enter with uiv.desktop.keyboard.press('Enter') afterwards.
Scrolling and holding keys
Available in extension 10.0.230+. These operations use Playwright's names and argument syntax.
Ui.Vision macro calls wait automatically, so omit async and await.
| Method | Behavior | Requirements |
|---|---|---|
uiv.browser.mouse.wheel(deltaX, deltaY) |
Trusted scrolling at the current pointer in viewport CSS pixels. Positive X goes right; positive Y goes down. Move the pointer over a nested panel first with uiv.browser.mouse.move(x, y). Returns after dispatch; wait for the intended content change afterwards. |
Chrome, Brave, Edge or another Chromium browser with the debugger API. |
uiv.desktop.mouse.wheel(deltaX, deltaY) |
Native scrolling at the OS pointer. Positive X goes right; positive Y goes down. macOS sends pixel events, Wayland continuous deltas, Windows converts 100 nominal pixels per notch, and X11 accumulates whole notches. The app and OS determine the resulting distance. | Desktop Automation 2.1.11+. Works from the extension or standalone desktop app. |
match.scrollIntoViewIfNeeded({timeout}) |
Reveals an attached, visible, stable DOM target, including inside nested scroll containers. A fully visible target stays in place. Also available as uiv.page.scrollIntoViewIfNeeded(locator, options). Find the target again for fresh coordinates afterwards. |
DOM target; all supported browsers, including Firefox. |
uiv.browser.keyboard.down(key)uiv.browser.keyboard.up(key) |
Hold or release one key. Consecutive down calls hold multiple keys; repeated down sends repeat events. Names include a, KeyW, ArrowDown, Shift, and ControlOrMeta (Meta on Mac, Control elsewhere). Hold modifiers separately; these methods take one key per call. |
Chromium debugger API. The desktop equivalent is uiv.desktop.keyboard.down/up(key). |
scrollIntoViewIfNeeded takes a timeout in milliseconds, matching Playwright;
existing finder timeouts such as uiv.$(locator, {timeout: 3}) use seconds.
An omitted scroll timeout uses !TIMEOUT_WAIT converted to milliseconds;
0 disables it. Oversized targets are centered as far as their containers allow.
The standalone desktop app has no DOM or browser tier. Legacy XMouseWheel
retains its existing notch units and axis mapping.
const panel = uiv.$('#panel');
uiv.browser.mouse.move(panel.x, panel.y);
uiv.browser.mouse.wheel(0, 300);
uiv.$('#target', {scroll: false}).scrollIntoViewIfNeeded({timeout: 3000});
const keyboard = uiv.browser.keyboard;
try {
keyboard.down('a');
keyboard.down('w');
uiv.sleep(1000);
} finally {
keyboard.up('w');
keyboard.up('a');
}
Pair key presses with releases in try/finally. The runner also releases held CDP keys
on completion, failure, or Stop. Held modifiers accompany browser mouse and wheel events.
Restore JavaScript demos in Settings for Core/DemoScrolling.js and
Core/DemoHeldKeys.js. Set NATIVE=true in DemoScrolling to test the native wheel.
Page, Reading & OCR
| Method | What it does | Example |
|---|---|---|
uiv.goto(url) |
Navigate and wait for the page load | uiv.goto('https://ui.vision') |
uiv.evaluate(code) |
Run JavaScript inside the website and return the result (the code must use return) | uiv.evaluate('return document.title') |
match.text / match.value |
Every DOM match carries its text and value — no separate store commands | uiv.$('css=h1').text |
uiv.ocr.read(opts) |
OCR the viewport, a region ({area}), a saved screenshot ({image}) or the desktop ({scope}) | uiv.ocr.read({area: myMatch}) |
AI, Tabs, Data & Files
| Method | What it does | Example |
|---|---|---|
uiv.ai.ask(prompt, opts) |
One round trip to the configured LLM; {images: [...]} attaches screenshots, {json: true} returns parsed JSON | uiv.ai.ask('total?', {images: [uiv.shot.viewport()], json: true}) |
uiv.ai.computerUse(task) |
Hands the whole task to the AI computer-use agent; returns its report | uiv.ai.computerUse('fill in this form') |
uiv.tabs.select / open / close / list |
Tab control with absolute indexes (1..N); every call returns {index, title, url} | uiv.tabs.select(2) |
uiv.csv.read / write / append / list / exists |
CSV files as plain 2D arrays — no !csvLine, no line-number bookkeeping | uiv.csv.append('log.csv', [ts, value]) |
uiv.download(source, opts) |
Download via link locator, URL or trigger function; renames ({as}), waits for completion, returns the on-disk name | uiv.download('css=a.pdf', {as: 'report.pdf'}) |
uiv.shot.viewport / page / element / desktop / area |
Screenshots (page = whole page, scroll-stitched); each returns the file name, so shots pipe into OCR or AI | uiv.ocr.read({image: uiv.shot.page()}) |
uiv.exportToDownloads(name) |
Copy a .png, .csv or 'log' from Ui.Vision storage to the browser's Downloads folder | uiv.exportToDownloads('report.csv') |
Variables, Logging & Utilities
| Method | What it does | Example |
|---|---|---|
uiv.getVar / uiv.setVar |
Read/write the same variable pool the classic commands use — internal !-variables included | uiv.setVar('!TIMEOUT_WAIT', 20) |
uiv.log(text, color) |
Write to the log (green/red/blue/..., '#shownotification' shows a browser notification) | uiv.log('done', 'green') |
uiv.banner(html, opts) |
On-page overlay message for the person watching — survives navigation, click-through, great for attended runs and human hand-offs | uiv.banner('Your turn: solve the captcha') |
uiv.sleep(ms) |
Fixed wait — last resort; finders auto-wait, so most macros need none | uiv.sleep('2s') |
uiv.run(command, target, value) |
The legacy bridge: run any classic command from inside a script | uiv.run('setProxy', ...) |
// @include + uiv.main |
Include another script file as a library; uiv.main is true only in the file that was started | // @include Core/Sub/MyLib.js |
Error Handling Is Just JavaScript
A failed uiv call throws a real exception. try/catch handles retries and fallbacks,
throw new Error('...') fails the macro deliberately, and optional elements are a
{required: false} option away — no !errorignore, no !statusOK bookkeeping.
A script should prove its own success: end it with a finder on something unique to the goal state,
or compare a read value and throw on mismatch.
Where To Go Next
Selenium-derived commands vs uiv.* — the conversion table for everyone coming from classic table macros. Or just try it: the AI sidebar writes uiv.* scripts for you — and converts your existing classic macros on request. The preinstalled "Draw a cat🐱" macro is a fun working example of finders, trusted input and self-verification.
Anything wrong or missing on this page? Suggestions?
...then please contact us.