Skip to main content

Overview

Ec2Instance deploys EC2 capacity behind an Auto Scaling Group with a launch template, a security group, and a graceful-termination handler. Use it for workloads that need direct server access, a custom AMI, or persistent block storage. Defaults applied by the construct:

What this creates

Every Ec2Instance synthesises the following resources:
  • A launch template (<serviceName>LaunchTemplate) carrying the AMI, instance type, block devices, IMDSv2 settings, and security group.
  • An Auto Scaling Group with a 60-second cooldown, all group metrics enabled, and termination policies OLDEST_LAUNCH_CONFIGURATION then CLOSEST_TO_NEXT_INSTANCE_HOUR.
  • A security group (AsgSecurityGroup), unless you pass your own via securityGroup.
  • A graceful-termination handler: an SQS queue with a dead-letter queue, a Node.js 24 Lambda (300s timeout, 256 MB), an EventBridge rule, and an ASG EC2_INSTANCE_TERMINATING lifecycle hook with a 600-second heartbeat and a CONTINUE default result.
  • A custom resource that suspends the AZRebalance scaling process.
  • A VPC, only when you omit the vpc prop.
  • A key pair (<serviceName>KeyPair), only when sshAccess is set without a keyName.
  • A persistent EBS data volume plus re-attach Lambdas, only when persistentDataVolume is set.
The termination hook is why a rolling update takes longer than a bare ASG replacement: each instance drains before the ASG releases it.

Resource class

The deep lib/... path is a published subpath of @fjall/components-infrastructure and is how you use a construct class directly. The scaffolded fjall/<app>/infrastructure.ts never does this. It imports App and the *Factory symbols from the package root instead.

Basic usage

Configuration options

Core properties

Capacity

minCapacity and maxCapacity forward straight to the CDK Auto Scaling Group without a Fjall-supplied fallback. When both are omitted, CDK’s default of 1 applies to each. Synth fails when spotCapacityPercentage is not an integer between 0 and 100, when minCapacity is negative or fractional, or when maxCapacity is below minCapacity.

Instance

Empty-string tag keys or values are rejected at synth.

Network

Lifecycle and storage

A persistent data volume is AZ-local, so vpcSubnets.availabilityZones must hold exactly one entry and it must match persistentDataVolume.availabilityZone. Synth fails otherwise. Changing dataVolumeStableOwnerId after deploy strands the existing volume, because the lifecycle Lambdas filter on it as a tag.

Machine images

When machineImage is omitted, the construct uses MachineImage.latestAmazonLinux2023().

Amazon Linux 2023 (default)

Ubuntu

Match the AMI architecture to the instance family. t4g is Graviton, so pick the arm64 parameter.

Custom AMI

User data

Amazon Linux 2023 ships dnf. Swap the package manager to match a non-default AMI.

Basic script

Container host

Storage

blockDevices replaces the default mapping, it does not extend it. Pass an entry for the /dev/xvda root volume alongside any extra volumes, or the root falls back to the AMI’s own default and loses the wrapper-enforced encryption.

Additional EBS volumes

The construct builds its own default through the safeEbs() helper, which pins encrypted: true and gp3. Import it from @fjall/components-infrastructure/lib/resources/aws/compute/blockDeviceVolume if you prefer that shape over hand-written CDK options.

RAID 0 across two volumes

For storage that must survive an instance refresh, use persistentDataVolume rather than a block device. Block devices are recreated empty with every replacement.

Security

Session Manager is the default access path

Session Manager needs no inbound rule and no public IP, which makes it the access path to reach for. It has one prerequisite: pass a role. CDK attaches the instance profile only when a role exists, so an Ec2Instance without role launches with no instance profile and ssmSessionPermissions has nothing to act on. With a role supplied, ssmSessionPermissions defaults to true and adds AmazonSSMManagedInstanceCore to it. Reach an instance with:
Pass ssmSessionPermissions: false when your role already carries scoped Session Manager statements, because the managed policy grants account-wide ssm:GetParameter.

sshAccess

sshAccess opts the group into SSH. Each field is one decision, and the construct makes none of them for you: Omit sshAccess and the construct creates no key pair, adds no port-22 rule and keeps the private placement. That is the default, and Session Manager covers it. A bastion reachable from one office address:
A build runner that stays private, reached over the VPN with a key pair the team already holds:
When the construct creates the key pair, AWS stores the private key in Systems Manager Parameter Store as /ec2/keypair/<key-pair-id>. Read it once and keep it out of the repository:
"0.0.0.0/0" still deploys, because a world-open rule is sometimes the intent, but it synthesises with a warning that names the instance (ec2WorldOpenSsh). The warning reads the placement the group resolved to, not public alone: it says “the entire internet” when the instances are internet-facing, through public: true or through public vpcSubnets together with associatePublicIpAddress: true, and otherwise that the rule admits every address that can route to them. Scope allowedIpCidrs to the addresses that need access instead. Adding a narrower range beside it restricts nothing, because both rules land on the same security group.
Synth rejects, naming the reason:
  • sshAccess alongside securityGroup. The construct adds ingress only to a security group it owns. Drop securityGroup, or keep it and add the port-22 rule there yourself with no sshAccess.
  • An empty allowedIpCidrs, a duplicated range, a literal that is not dotted-quad/prefix-length, or a list token such as valueAsList or Fn.split over a parameter in place of the list, even spread into a literal one. A token per element passes through, and Fn.split over a literal string is already a list. IPv6 ranges are not accepted.
  • An empty keyName, or a keyName equal to <serviceName>KeyPair. Omit it to have a key pair created and kept.
  • public: true with associatePublicIpAddress: false, with a vpcSubnets.subnetType other than public, or with a vpcSubnets selector by group name or instance.
  • A field sshAccess does not have. The pre-34 allowedCidrs is allowedIpCidrs here.
  • null in place of the object. Omit sshAccess instead.
The retired enableSSH is still read at synth so an older file cannot deploy silently without SSH. enableSSH: true fails with the translation to sshAccess. enableSSH: false asked for the default and still gets it, with a warning to delete the prop (ec2RetiredSshProp). Changing keyName or public on a deployed group changes the launch template, so the next deploy runs an instance refresh under updatePolicy. Changing allowedIpCidrs edits the security group’s rules in place.

Bring your own security group

IAM role

The role prop takes a standard CDK IAM Role from aws-cdk-lib/aws-iam.

Scaling and rollout

Across availability zones

Spot capacity

spotCapacityPercentage sets the spot share above base capacity, using the price-capacity-optimized allocation strategy. Any value above 0 switches the ASG to a mixed instances policy.

Update policy

Any change to the launch template (user data, AMI, block devices) triggers a rollout. updatePolicy picks the strategy. Default healthy percentages are 100 / 200, which surges then shrinks and avoids a capacity dip. With persistentDataVolume set they become 0 / 100, because a single-attach EBS volume cannot serve a surge instance.

Properties and methods

Patterns

Web fleet behind a load balancer

Operations host

Reach it over Session Manager rather than SSH, so it needs no inbound rule and no public IP. The role is what makes that work.
The tags entry propagates to launched instances, so SSM SendCommand can target tag:Role = ops-host. The instance also needs a route to the SSM endpoints, either NAT egress from a private subnet or VPC interface endpoints for ssm, ssmmessages, and ec2messages.

GPU instance

Complete example

Best practices

  1. Reach instances through Session Manager. When someone must have SSH, scope sshAccess.allowedIpCidrs to their addresses and leave public off unless they arrive from the internet.
  2. Keep the encrypted /dev/xvda root entry whenever you override blockDevices.
  3. Supply a scoped securityGroup rather than mutating the generated one after construction.
  4. Always pass a role. Without one there is no instance profile, so Session Manager, S3 reads, and every other AWS call from the instance fail. Pass ssmSessionPermissions: false when that role already carries scoped Session Manager statements.
  5. Use persistentDataVolume for state that must survive an instance refresh, and pin dataVolumeStableOwnerId for the life of the stack.
  6. Set tags so SSM SendCommand and cost allocation can target the fleet.
  7. Pick the AMI architecture that matches the instance family. Graviton families (t4g, m6g, c7g) need arm64 images.

Cost optimisation

  • Set spotCapacityPercentage for fault-tolerant workloads, with capacityRebalance: true.
  • Pass instanceMonitoring: Monitoring.BASIC to drop from 1-minute to 5-minute metrics. Detailed monitoring is on by default and is billed per instance.
  • Right-size instanceType from CloudWatch metrics, and prefer Graviton families where the workload supports arm64.
  • Use gp3 volumes rather than gp2, which the construct default already does.
  • Set maxInstanceLifetime to recycle long-lived instances instead of running oversized ones indefinitely.
  • Consider Savings Plans for steady-state capacity.

Next Steps

Compute Factory

Build EC2 and ECS compute through the Fjall compute factory pattern.

Security Group

Control inbound and outbound traffic for your instances.

VPC

Configure the network your EC2 instances run in.

IAM Role

Grant least-privilege AWS permissions to your instances.