Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
82 changes: 74 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -93,6 +93,11 @@ crawlscope validate --url https://example.com --sitemap https://example.com/site

Child sitemap indexes are supported automatically.

Set `CRAWLSCOPE_PROFILE_TOKEN` through the process environment or your secret
manager when the site protects diagnostic timing with `X-Profile-Token`.
Crawlscope intentionally has no `--profile-token` flag because command arguments
can be exposed through shell history and process listings.

Validation output is grouped for terminal scanning:

```text
Expand Down Expand Up @@ -141,10 +146,69 @@ puts result.issues.to_a.map(&:message)
- `urls`: sitemap URLs selected for validation
- `pages`: fetched page snapshots
- `issues`: structured issues with `code`, `severity`, `category`, `url`, and `message`
- `server_timing_summary`: aggregate response timing data when pages publish it

`result.ok?` returns `false` when an error is present. Warnings and notices
remain available through `result.issues` without making the result fail.

## Server Timing

Crawlscope parses the
[`Server-Timing`](https://www.w3.org/TR/server-timing/) response header for both
HTTP and browser-rendered crawls. Each page exposes its parsed header through
`page.server_timing`:

```ruby
page.server_timing.each do |metric|
puts [metric.name, metric.duration, metric.description].compact.join(": ")
end
```

The validation report adds a `Server Timing` section only when at least one page
publishes the header. It includes:

- header coverage and the number of pages publishing durations
- sample and page counts with average, p50, p95, and maximum per metric
- non-duration signals, including cache status and routing descriptions
- the ten pages with the largest individual metric
- the number of malformed entries ignored during parsing

`dur` values are reported as milliseconds, as recommended by the specification.
Crawlscope does not add durations together because metrics such as `total`,
`app`, and `db` may overlap. A page's worst-offender entry is its largest
individual metric instead.

The current HTTP and browser transports expose response headers, not response
trailers. Metrics published only in a `Server-Timing` trailer are therefore not
available to Crawlscope.

Rails 8 applications can enable `ActionDispatch::ServerTiming` with:

```ruby
config.server_timing = true
```

Applications that expose timing only to authenticated diagnostics can set
`config.profile_token`. When this value is present, Crawlscope sends it as
`X-Profile-Token` on same-origin sitemap, HTTP, and browser document requests.
The header survives same-origin redirects but is removed before a cross-origin
redirect, and browser subresources never receive it. Keep the token in the host
application's secret store rather than a URL, command argument, report, or
checked-in configuration.

`CRAWLSCOPE_PROFILE_TOKEN` is the portable default for Rails, the standalone
CLI, Rake tasks, and plain Ruby callers. Rails applications may instead assign
`config.profile_token` from encrypted credentials when that is their established
secret-management boundary; explicit configuration takes precedence over the
environment.

New Rails applications enable it in development by default; production remains
opt-in. Rails publishes the Active Support notification names observed during
each request and sums repeated events with the same name. The exact metrics
therefore depend on the application and request, but commonly include controller,
view, database, and cache instrumentation. Crawlscope accepts Rails' dotted
metric names and reports each metric independently.

## Rails Usage

Run the install generator after adding the gem:
Expand All @@ -162,10 +226,11 @@ Customize the `Crawlscope.configure` block inside the generated initializer:

```ruby
Crawlscope.configure do |config|
config.base_url = -> { "https://example.com" }
config.sitemap_path = -> { Rails.public_path.join("sitemap.xml").to_s }
config.site_name = "Example"
config.schema_registry = -> { Crawlscope::SchemaRegistry.default }
config.base_url = -> { ENV.fetch("CRAWLSCOPE_BASE_URL", "http://localhost:3000") }
config.sitemap_path = lambda {
ENV.fetch("SITEMAP", "#{config.base_url.to_s.chomp("/")}/sitemap.xml")
}
config.site_name = ENV.fetch("CRAWLSCOPE_SITE_NAME", "Application")
end
```

Expand Down Expand Up @@ -193,6 +258,7 @@ Available environment overrides:

- `URL`
- `SITEMAP`
- `CRAWLSCOPE_PROFILE_TOKEN`
- `RULES=metadata,links`
- `JS=1` or `RENDERER=browser`
- `TIMEOUT=30`
Expand Down Expand Up @@ -226,10 +292,10 @@ bundle exec rake 'crawlscope:validate:ldjson[https://example.com/article]'

`crawlscope:validate` runs all default sitemap rules: indexability, metadata,
structured data, uniqueness, content quality, and links. `URL` is the site
base. Without `SITEMAP`, Crawlscope uses the configured sitemap path, then
falls back to `/sitemap.xml`. With `SITEMAP`, Crawlscope uses `URL` as the site
base and validates URLs from that sitemap. `SITEMAP` may be a full URL or a
local file path.
base. Without `SITEMAP`, Crawlscope uses the configured sitemap URL, then fetches
`/sitemap.xml` from `URL` over HTTP. With `SITEMAP`, Crawlscope uses `URL` as
the site base and validates URLs from that sitemap. `SITEMAP` may be a full URL
or an explicitly selected local file path.

Plain `rake` does not pass `--url` style flags to tasks. Use `URL=...` or the
task-argument form above instead.
Expand Down
28 changes: 28 additions & 0 deletions UPGRADE.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,34 @@ behavior.

## Next Release

### Default sitemap is fetched over HTTP

Crawlscope no longer prefers `public/sitemap.xml` when validating a localhost
application. Without an explicit `SITEMAP` or configured `sitemap_path`, it now
fetches `/sitemap.xml` from the configured base URL so the crawl observes the
live application and database state.

The Rails installer now generates the same HTTP default. Regenerate or update
existing initializers that use `Rails.public_path.join("sitemap.xml")`.
Explicit local sitemap paths remain supported through `SITEMAP`, `--sitemap`,
or `config.sitemap_path`.

`CRAWLSCOPE_PROFILE_TOKEN` is now the portable default for standalone CLI, Rake,
Rails, and plain Ruby usage. Explicit `config.profile_token` values still take
precedence.

### Server-Timing report output

No host configuration is required. When one or more responses publish a
`Server-Timing` header, the text report now includes an optional `Server Timing`
section. Host applications that parse report text should accept this additional
section.

Parsed metrics are available through `page.server_timing`, and aggregate data is
available through `result.server_timing_summary`. Crawlscope interprets `dur`
values as milliseconds and ignores malformed entries while reporting their
count.

### Ruby 3.3 is now required

Crawlscope now depends on the current Async runtime for production async HTTP
Expand Down
19 changes: 18 additions & 1 deletion lib/crawlscope/browser.rb
Original file line number Diff line number Diff line change
Expand Up @@ -4,13 +4,15 @@

module Crawlscope
class Browser
def initialize(base_url:, timeout_seconds:, network_idle_timeout_seconds:, scroll_page:)
def initialize(base_url:, timeout_seconds:, network_idle_timeout_seconds:, scroll_page:, profile_token: nil)
@base_url = base_url
@timeout_seconds = timeout_seconds
@network_idle_timeout_seconds = network_idle_timeout_seconds
@profile_token = profile_token
@scroll_page = scroll_page
@browser = build_browser
@page = @browser.create_page
configure_profile_requests
end

def close
Expand Down Expand Up @@ -72,6 +74,21 @@ def build_browser
)
end

def configure_profile_requests
return if @profile_token.to_s.empty?

@page.network.intercept(resource_type: :document)
@page.on(:request) do |request|
headers = RequestHeaders.add_profile_token(
request.headers.dup,
url: request.url,
base_url: @base_url,
profile_token: @profile_token
)
request.continue(headers: headers.map { |name, value| {name: name.to_s, value: value.to_s} })
end
end

def scroll_for_render
@page.evaluate("(function() { if (document.body) { window.scrollTo(0, document.body.scrollHeight); } })()")
wait_for_network_idle
Expand Down
7 changes: 6 additions & 1 deletion lib/crawlscope/configuration.rb
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ class Configuration
RENDERERS = %i[http browser].freeze
DEFAULT_TIMEOUT_SECONDS = 20

attr_writer :allowed_statuses, :base_url, :browser_factory, :concurrency, :fetch_executor, :network_idle_timeout_seconds, :output, :renderer, :rule_registry, :schema_registry, :scroll_page, :site_name, :sitemap_path, :timeout_seconds
attr_writer :allowed_statuses, :base_url, :browser_factory, :concurrency, :fetch_executor, :network_idle_timeout_seconds, :output, :profile_token, :renderer, :rule_registry, :schema_registry, :scroll_page, :site_name, :sitemap_path, :timeout_seconds

def allowed_statuses
value = resolve(@allowed_statuses)
Expand Down Expand Up @@ -59,6 +59,10 @@ def output
value.nil? ? $stdout : value
end

def profile_token
resolve(@profile_token) || ENV["CRAWLSCOPE_PROFILE_TOKEN"]
end

def renderer
value = resolve(@renderer)
normalized_value = value.to_s.strip
Expand Down Expand Up @@ -93,6 +97,7 @@ def audit(base_url: self.base_url, sitemap_path: self.sitemap_path, rule_names:
concurrency: concurrency,
fetch_executor: fetch_executor,
network_idle_timeout_seconds: network_idle_timeout_seconds,
profile_token: profile_token,
renderer: renderer,
timeout_seconds: timeout_seconds,
allowed_statuses: allowed_statuses,
Expand Down
7 changes: 5 additions & 2 deletions lib/crawlscope/crawl.rb
Original file line number Diff line number Diff line change
Expand Up @@ -2,14 +2,15 @@

module Crawlscope
class Crawl
def initialize(base_url:, sitemap_path:, rules:, schema_registry:, browser_factory: nil, concurrency: Configuration::DEFAULT_CONCURRENCY, fetch_executor: nil, network_idle_timeout_seconds: Configuration::DEFAULT_BROWSER_NETWORK_IDLE_TIMEOUT_SECONDS, renderer: :http, scroll_page: Configuration::DEFAULT_BROWSER_SCROLL_PAGE, timeout_seconds: Configuration::DEFAULT_TIMEOUT_SECONDS, allowed_statuses: Configuration::DEFAULT_ALLOWED_STATUSES)
def initialize(base_url:, sitemap_path:, rules:, schema_registry:, browser_factory: nil, concurrency: Configuration::DEFAULT_CONCURRENCY, fetch_executor: nil, network_idle_timeout_seconds: Configuration::DEFAULT_BROWSER_NETWORK_IDLE_TIMEOUT_SECONDS, profile_token: nil, renderer: :http, scroll_page: Configuration::DEFAULT_BROWSER_SCROLL_PAGE, timeout_seconds: Configuration::DEFAULT_TIMEOUT_SECONDS, allowed_statuses: Configuration::DEFAULT_ALLOWED_STATUSES)
@base_url = base_url
@sitemap_path = sitemap_path
@rules = Array(rules)
@schema_registry = schema_registry
@browser_factory = browser_factory
@concurrency = concurrency
@network_idle_timeout_seconds = network_idle_timeout_seconds
@profile_token = profile_token
@renderer = renderer.to_sym
@fetch_executor = fetch_executor || default_fetch_executor
@scroll_page = scroll_page
Expand Down Expand Up @@ -49,6 +50,7 @@ def sitemap_urls
adapter: http_adapter,
concurrency: @concurrency,
fetch_executor: @fetch_executor,
profile_token: @profile_token,
timeout_seconds: @timeout_seconds
).urls(base_url: @base_url)
raise ValidationError, "No URLs found in sitemap: #{@sitemap_path}" if urls.empty?
Expand All @@ -61,6 +63,7 @@ def browser
base_url: @base_url,
timeout_seconds: @timeout_seconds,
network_idle_timeout_seconds: @network_idle_timeout_seconds,
profile_token: @profile_token,
scroll_page: @scroll_page
)
rescue LoadError => error
Expand All @@ -71,7 +74,7 @@ def page
if @renderer == :browser
(@browser_factory || method(:browser)).call
else
Http.new(base_url: @base_url, timeout_seconds: @timeout_seconds, adapter: http_adapter)
Http.new(base_url: @base_url, timeout_seconds: @timeout_seconds, adapter: http_adapter, profile_token: @profile_token)
end
end

Expand Down
15 changes: 13 additions & 2 deletions lib/crawlscope/http.rb
Original file line number Diff line number Diff line change
Expand Up @@ -10,10 +10,11 @@ class Http
MAX_REDIRECTS = 5
USER_AGENT = "Mozilla/5.0 (compatible; Crawlscope/1.0)"

def initialize(base_url:, timeout_seconds:, adapter: nil)
def initialize(base_url:, timeout_seconds:, adapter: nil, profile_token: nil)
@base_url = base_url
@timeout_seconds = timeout_seconds
@adapter = adapter
@profile_token = profile_token
@connections_by_thread = Concurrent::Map.new
end

Expand All @@ -28,6 +29,12 @@ def close
def fetch(url)
response = connection.get(url) do |request|
request.headers["User-Agent"] = USER_AGENT
RequestHeaders.add_profile_token(
request.headers,
url: url,
base_url: @base_url,
profile_token: @profile_token
)
end

final_url = response.env.url.to_s
Expand Down Expand Up @@ -67,7 +74,11 @@ def fetch(url)
def connection
@connections_by_thread.compute_if_absent(Thread.current.object_id) do
Faraday.new do |faraday|
faraday.response :follow_redirects, limit: MAX_REDIRECTS
faraday.response(
:follow_redirects,
limit: MAX_REDIRECTS,
callback: RequestHeaders.method(:strip_profile_token_on_cross_origin_redirect)
)
faraday.options.timeout = @timeout_seconds
faraday.options.open_timeout = @timeout_seconds
faraday.adapter @adapter if @adapter
Expand Down
8 changes: 8 additions & 0 deletions lib/crawlscope/page.rb
Original file line number Diff line number Diff line change
Expand Up @@ -19,5 +19,13 @@ def initialize(url:, normalized_url:, final_url:, normalized_final_url:, status:
def html?
!doc.nil?
end

def header(name)
headers.find { |key, _value| key.to_s.casecmp?(name.to_s) }&.last
end

def server_timing
@server_timing ||= ServerTiming.new(header("server-timing"))
end
end
end
13 changes: 9 additions & 4 deletions lib/crawlscope/reporter.rb
Original file line number Diff line number Diff line change
Expand Up @@ -19,13 +19,18 @@ def report(result)

if result.issues.size.zero?
@io.puts("Status: OK")
return
else
@io.puts("Status: #{status_for(result.issues)}")
@io.puts("Issues: #{result.issues.size} total (#{severity_summary(result.issues)})")
end

@io.puts("Status: #{status_for(result.issues)}")
@io.puts("Issues: #{result.issues.size} total (#{severity_summary(result.issues)})")
@io.puts("")
ServerTiming::Reporter.new(io: @io).report(
result.server_timing_summary,
base_url: result.base_url
)
return if result.issues.size.zero?

@io.puts("")
report_summary(result.issues)
@io.puts("")
report_issue_groups(result.issues, base_url: result.base_url)
Expand Down
34 changes: 34 additions & 0 deletions lib/crawlscope/request_headers.rb
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
# frozen_string_literal: true

require "uri"

module Crawlscope
module RequestHeaders
PROFILE_TOKEN = "X-Profile-Token"

module_function

def add_profile_token(headers, url:, base_url:, profile_token:)
return headers if profile_token.to_s.empty?
return headers unless same_origin?(url, base_url: base_url)

headers[PROFILE_TOKEN] = profile_token
headers
end

def strip_profile_token_on_cross_origin_redirect(response_env, request_env)
return if same_origin?(request_env[:url], base_url: response_env[:url])

request_env[:request_headers].delete(PROFILE_TOKEN)
end

def same_origin?(url, base_url:)
url_uri = URI.join(base_url.to_s, url.to_s)
base_uri = URI.parse(base_url.to_s)

[url_uri.scheme, url_uri.host, url_uri.port] == [base_uri.scheme, base_uri.host, base_uri.port]
rescue URI::Error
false
end
end
end
4 changes: 4 additions & 0 deletions lib/crawlscope/result.rb
Original file line number Diff line number Diff line change
Expand Up @@ -5,5 +5,9 @@ module Crawlscope
def ok?
issues.none?(&:error?)
end

def server_timing_summary
ServerTiming::Summary.new(pages)
end
end
end
Loading
Loading