May 29, 20269 min read

Validating and sanitizing user input with Yup and DOMPurify
February 17, 2021 · 5 min read
A working set of Yup schemas I keep reaching for, and the DOMPurify pattern that catches what validation can’t.
Most of the time when a form blows up in production, the cause is something I should have caught at the boundary. A phone number with a typo, an email that’s just “asdf”, an HTML comment a user pasted in that broke the markup downstream. Each one is the kind of thing I should have validated, didn’t, and paid for later.
This post is the set of schemas and helpers I keep reusing across projects, plus the DOMPurify pattern that handles the output side. Validation makes sure the data coming in matches what I expect; sanitization makes sure the data going back out can’t run as code in someone else’s browser. They solve different problems and you usually want both.

Validation vs. sanitization
Quick definition pass before code, since these two get conflated all the time.
- Validation is checking that data fits the rules before you accept it. Is this an email? Does this username only contain the characters I allow? Is this number positive? Validate on the client for UX, validate on the server because the client is untrusted.
- Sanitization is removing dangerous content before you render it. The data might be technically valid (a string is a string), but if it contains a
<script>tag and you echo it back into the page, you’ve shipped an XSS vulnerability.
Yup handles the first job. DOMPurify handles the second.
Yup basics
Yup is a schema library: you describe what valid data looks like, then call .validate() on the input. It throws if the data doesn’t match. It also chains, which makes the rules read top-to-bottom.
Here are the schemas I reuse most often.
import * as yup from 'yup';
const emailSchema = yup.string().email('Must be a valid email address').required('Email is required');
Yup’s built-in email check is good enough for most use cases. If you need stricter rules (e.g. blocking disposable email domains), chain a .test() on the end with your own predicate.
Phone number (digits only)
const phoneSchema = yup.string()
.matches(/^[0-9]{10,15}$/, 'Must be a valid phone number with 10 to 15 digits')
.required('Phone number is required');
This deliberately doesn’t try to handle international formatting (parentheses, dashes, country codes). For that I reach for libphonenumber-js instead and run Yup’s .test() against its isValidPhoneNumber() result. Trying to do international phone validation in regex is a road I don’t recommend.
Git commit hash
const gitHashSchema = yup.string()
.matches(/^[0-9a-f]{40}$/, 'Must be a 40-character hex string')
.required('Commit hash is required');
Useful when building tooling around code. If I want to also accept short SHAs, I loosen the bound to {7,40}.
IPv4 address
const ipSchema = yup.string()
.matches(/^(25[0-5]|2[0-4]\d|[01]?\d?\d)(\.(25[0-5]|2[0-4]\d|[01]?\d?\d)){3}$/, 'Must be a valid IPv4 address')
.required('IP address is required');
The regex is ugly but correct: each octet is 0-255. For IPv6 you’ll want a different pattern (or a library), but for IPv4 this holds up.
Ethereum wallet address
const ethWalletSchema = yup.string()
.matches(/^0x[a-fA-F0-9]{40}$/, 'Must be a valid Ethereum wallet address')
.required('Wallet address is required');
An Ethereum address is 0x followed by 40 hex characters. This regex catches the format but not the checksum (EIP-55). If you need full checksum validation, use ethers.utils.isAddress() from ethers.js inside a .test().
Object schemas: validating a whole form
Single-field schemas are useful, but most real validation runs over a whole object. yup.object().shape({...}) lets you compose field rules and validate them as a unit.
const userSchema = yup.object().shape({
username: yup.string()
.matches(/^[a-zA-Z0-9_]+$/, 'Username can only contain letters, numbers, and underscores')
.min(3, 'Username must be at least 3 characters long')
.max(20, 'Username cannot be longer than 20 characters')
.required('Username is required'),
password: yup.string()
.min(8, 'Password must be at least 8 characters long')
.required('Password is required'),
email: yup.string()
.email('Please provide a valid email')
.required('Email is required'),
});
const userData = {
username: 'John_Doe',
password: 'securePa55',
email: 'john.doe@example.com',
};
userSchema.validate(userData)
.then(validData => console.log('Valid user data:', validData))
.catch(err => console.error('Validation error:', err.errors));
By default, .validate() stops at the first error. Pass { abortEarly: false } if you want every error at once, which is what most form libraries (e.g. Formik, React Hook Form with the Yup resolver) want so they can show all the field errors together.
DOMPurify: cleaning HTML on the way out
Validation tells you the data is shaped right. It doesn’t help when the data is a string that, if rendered raw into the page, will execute. That’s where DOMPurify comes in: you hand it untrusted HTML, it hands you back HTML with the dangerous bits removed.
Browser usage:
import DOMPurify from 'dompurify';
const userInputHTML = `
<div><strong>Hello</strong> world!</div>
<img src="http://example.com/image.jpg" onerror="alert('XSS')">
<script>alert('HACKED!');</script>
`;
const sanitizedOutput = DOMPurify.sanitize(userInputHTML);
// sanitizedOutput now has the script tag and onerror attribute removed.
The <script> tag is gone. The onerror attribute on the image is gone. The benign <strong> and <div> survive. That’s the right behavior for rendering user-generated content like comments or descriptions.
For Node.js you need a DOM to give DOMPurify something to operate on. jsdom handles that:
npm install dompurify jsdom
import { JSDOM } from 'jsdom';
import createDOMPurify from 'dompurify';
const { window } = new JSDOM('');
const DOMPurify = createDOMPurify(window);
const maliciousHTML = `
<h1>Welcome</h1>
<script>alert('HACKED!');</script>
`;
const clean = DOMPurify.sanitize(maliciousHTML);
// 'clean' now contains safe HTML.
If you need to allow only a specific subset of tags (say, just <b>, <i>, and <a>), pass { ALLOWED_TAGS: [...] } as the second argument. The default config is sensible for most cases, but locking the allowlist down further is good practice when the input is something like a comment field.
What I actually do
The pattern I land on most projects:
- Yup schema for every form, with the resolver wired into whatever form library I’m using.
- The same schema runs again on the server before anything touches the database. The client schema is for UX; the server schema is the actual enforcement.
- Anywhere I render user-generated HTML back to the page, it goes through DOMPurify first. If the field doesn’t need to allow HTML at all, I escape instead and skip DOMPurify entirely.
If you want more Yup schema patterns – dates, nested objects, arrays of objects, conditional fields – I covered a wider set in a follow-up post. Drop a note if there’s a specific case you’d like to see.
Discussion
Have thoughts? Drop them in.
Comments are powered by Disqus. Sign in once, comment anywhere.

