TL;DR
- PostgreSQL with pgvector is our default for a GTM agent: durable records, constraints and retrieval can share one database.
- Redis is useful for temporary sessions, caches and coordination. Configure persistence deliberately; do not make an expiring cache the only copy of approval or suppression decisions.
- Qdrant is the specialised vector option for teams that want payload filtering and an open-source deployment route.
- Pinecone suits managed retrieval with namespace-based tenant separation. Standard has a $50 monthly usage minimum; the newer $20 Builder plan is a different feature purchase.
- Weaviate suits hybrid retrieval and explicit tenant collections. Flex starts at $45/month; the free cloud tier permits only three tenants.
- Separate business truth from retrieved context. A similar old note must not overwrite a current CRM stage or authorise a new send.
What a persistent agent needs to remember
An agent that researches an account today should be able to resume tomorrow without asking for the same records again. Persistence helps. The design goes wrong when all remembered material becomes equally authoritative.
A CRM opportunity stage is a business field. A website quotation is evidence. A summary is a derived interpretation. A worker's current tool result is temporary state. Storing all four as anonymous embeddings makes them easy to retrieve and hard to trust.
| Kind of memory | Example | Suitable home | How it changes |
|---|---|---|---|
| Durable business state | Account ID, suppression, accepted approval | Relational records with constraints | Explicit authorised update |
| Event history | Research requested; CRM update attempted | Durable event/run tables | Add event and record result |
| Retrieved evidence | Company page with source URL and date | Documents plus vector/text index | Refresh source and rebuild derived chunks |
| Working state | Short-lived session or cached API response | Redis or bounded runtime state | Expire or regenerate |
| Summary | Suggested account-fit explanation | Versioned derived record | Regenerate from current evidence |
This guide compares storage infrastructure for a persistent GTM agent. Our guide to running an agency on Claude Code addresses the operating approach around it, while n8n vs Zapier vs Make covers orchestration. A database does not itself provide the whole research or approval workflow.
The product facts and public prices below were assessed on 1 October 2026. Architecture and workload figures are editorial examples; no comparative latency, retrieval-quality or live database test is claimed.
Five storage options compared
| Option | Best job | Tenant mechanism | Cost basis | Main reason to choose something else |
|---|---|---|---|---|
| PostgreSQL + pgvector | Durable records and moderate retrieval together | Roles and row-level policies; explicit tenant keys | Self-hosting or managed database compute/storage | Retrieval needs independent scaling or specialist operations |
| Redis | Cache, sessions and coordination | Application keys plus appropriate access boundaries | Memory/resource tier and deployment | Need authoritative relational records and loss-resistant decisions |
| Qdrant | Specialised filtered vector retrieval | Payload tenant filters, sharding options | Self-host resources or dedicated cloud cluster | Need one database for transactions and retrieval |
| Pinecone | Provider-operated vector retrieval | Namespace per tenant in a serverless index | Plan floor and metered storage/read/write consumption | Want self-hosted control or immediate query freshness |
| Weaviate | Hybrid object/vector retrieval | Enabled multi-tenancy with tenant-specific shards | Cloud resources/dimensions/storage or self-hosting | Need transactional CRM state rather than retrieval |
1. PostgreSQL with pgvector: the default durable foundation
PostgreSQL provides relational records and transactions. You can commit a research decision and its pending delivery event together, rather than writing one and losing the other during a crash. Unique constraints can give an account or action a stable identity.
pgvector adds vector similarity search, including exact search and approximate HNSW/IVFFlat indexes. Keep an embedding beside its document ID, client ID and source version. For many early GTM agents, that avoids a second database, replication pipeline and deletion process.
There is an important retrieval detail: filtering an approximate index can return fewer candidates than requested. The Supabase pgvector documentation explains this and points to iterative scanning. A query that returns only two authorised chunks should not be “fixed” by dropping the client filter. Adjust retrieval/indexing or use an exact baseline for the bounded dataset.
PostgreSQL row-level security can constrain which rows normal database roles read or modify. Owners, superusers and BYPASSRLS roles have different behaviour. A policy does not isolate clients if the application runs every request through a privileged role that bypasses it.
There is no software subscription for PostgreSQL and pgvector themselves. For a managed example, Supabase Pro starts at $25/month and includes a $10 compute credit covering one Micro instance. Its page lists 8 GB disk per project, then $0.125/GB, and seven days of daily backups. Medium compute is $60/month before the credit; point-in-time recovery starts separately at $100/month. Neon's public pricing gives another basis: Launch compute is $0.106/CU-hour and database storage $0.35/GB-month, with history storage charged separately.
Choose this combination when the team needs reliable account state, approval history and a manageable evidence corpus. It is a poor choice only if the database is being asked to meet a retrieval load or operational requirement it cannot handle economically. Splitting the vector service should solve a measured problem, not follow an architecture diagram's fashion.
2. Redis: fast working state with explicit loss tolerance
Redis is useful for an agent session, a cached research response, rate-limit counters or coordination between workers. Expiry helps keep short-lived state bounded. It is also capable of richer data and search work, but this guide recommends it first as an optional working-state layer.
The persistence documentation distinguishes periodic RDB snapshots from append-only logging. An every-second AOF policy has a different loss window from a snapshot taken several minutes earlier. “Redis persists data” is therefore not a sufficient design statement. Persistence, replication and eviction settings all affect the role you can assign it.
The cloud pricing page lists Free at 30 MB, Essentials from $0.007/hour, presented as $5/month, and Pro with a $200/month minimum. Memory/resource size and topology change the purchase; $5 is an entry figure, not an arbitrary-size production cluster. Current Redis 8 licensing offers AGPLv3, RSALv2 and SSPLv1 choices; proprietary Redis products use separate commercial terms. Avoid applying an old BSD description to every current edition.
For a small research agent, cache content can be rebuilt. Suppression and accepted approval should have a durable authoritative copy. A lock expiry also does not prove that an external email was never sent: a delayed worker may finish after its lease expires. Use action IDs and provider-side deduplication where supported, alongside the durable action record.
Choose Redis when temporary-state speed or coordination is demonstrably useful. Skip it initially if PostgreSQL and the job runner already handle the modest workload. A second state store adds another place to investigate a stalled run.
3. Qdrant: specialised retrieval with deployment choice
Qdrant is an Apache-2.0 vector database that stores points with vectors and payload metadata. It offers a self-hosted service and managed cloud. This is useful when retrieval deserves its own operational boundary while the account and action tables remain in PostgreSQL.
The multitenancy guide describes payload partitioning, user-defined sharding and tiered multitenancy. For many small clients, a shared collection with an indexed tenant field and mandatory filtering is practical. Larger clients can justify different sharding. Those mechanisms do not authorise arbitrary callers: the application must derive the permitted tenant from the authenticated request and bind every retrieval to it.
The pricing page lists a free cloud tier with 1 GB RAM and 4 GB disk, intended for prototypes without high availability. Standard uses resource-based pricing with dedicated resources; Premium adds commercial security/support scope and a minimum spend. Hybrid Cloud runs managed clusters on your infrastructure, while Private Cloud is a distinct isolated deployment.
No universal dollar total follows from a vector count alone. Price the chosen memory, disk, region and replication topology. For self-hosting, include the node and maintenance; for managed Standard, use the cluster rate and selected support/backup scope. This basis is sufficient to shortlist Qdrant without inventing a $9 production package.
Choose Qdrant when filtered retrieval needs independent scaling or the team values a portable self-hosting route. It is a poor substitute for the relational approval and action ledger. Keeping those records in PostgreSQL also means the retrieval index can be rebuilt after a restore or embedding change.
4. Pinecone: managed retrieval, with plan and freshness distinctions
Pinecone suits a team that wants the provider to operate vector infrastructure. Its serverless multitenancy guidance uses a namespace per tenant, with data operations directed to that namespace. The application still decides which namespace the requesting user may access.
The current pricing page lists Builder at $20/month flat, Standard with a $50/month usage minimum, and Enterprise with a $500/month usage minimum. Standard adds features including backup/restore and RBAC/SSO; usage above the minimum is metered. Embeddings, reranking and Assistant services are distinct consumption items. A $50 floor is not an unlimited plan, and it is no longer accurate to describe every paid entry as $50.
Pinecone documents eventual consistency: a successful write can precede its visibility to a query. Its freshness guidance describes checking serverless log sequence numbers. For a GTM agent, do not put the only suppression check in an index that can momentarily return old content. Check the authoritative record before the consequential action.
Choose Pinecone when managed retrieval and its deployment choices reduce engineering work enough to justify the service. Standard is the relevant comparison when backup/restore and team controls are required; Builder is worth considering for a smaller application with different needs. Keep source documents and chunk IDs outside the index so migration does not depend on reconstructing business truth from embeddings.
5. Weaviate: hybrid retrieval with explicit tenants
Weaviate stores objects and vectors and supports combining structured filtering with retrieval. It is relevant where keyword and semantic search both matter, such as finding a precise product name within a broader account-research question.
Its multi-tenancy documentation describes enabling the feature on a collection and storing each tenant on a separate shard. Tenant-specific operations must select the intended tenant. This is a clearer storage structure than treating a free-text client name in a document as isolation.
The current cloud pricing lists a free tier with 100,000 objects, 1 GB memory, 10 GB disk, one collection and up to three tenants. Flex starts at $45/month, with paid capacity determined by the configuration, dimensions and storage. Premium uses prepaid commercial scope. Cloud embeddings and Query Agent services add their own consumption basis; a listed feature is not a promise of unlimited free model calls.
For a ten-client research agent, the free cloud tier's three-tenant limit already makes it a poor equivalent to Flex. The 100,000-object headline does not override that boundary or establish that a particular embedding/index configuration fits its memory.
Choose Weaviate when hybrid retrieval and tenant collection operations match the application. Keep CRM stages, approvals and action identity in the relational store. Adding an agent-memory service around a database may simplify application code, but it does not automatically establish the correct business authority or retention policy.
An architecture for a ten-client research agent
Consider an agency researching 1,000 accounts across ten clients. Two clients both have an account named Northstar. Their contacts, research sources and approved messages must stay distinct.
Use the CRM as the authority for its existing business fields. In PostgreSQL, keep local account references keyed by client ID and CRM account ID, not by display name. Add document records with source URL, retrieved timestamp and content hash. Chunks reference that document version and carry the client ID plus embedding-model version.
The run table records the account, requested research and current status. A proposed message has its own version and approval state. On approval, a transaction records the approved version and an outbox action with a unique client/action ID. A worker claims the pending action, checks the current approval and suppression record, performs the external operation and records the result.
That pattern handles a database crash between decision and queueing. It does not make an external send exactly once. If the provider accepts a message but the worker crashes before recording success, retrying can duplicate it. Use a supported external idempotency key or reconcile the provider result before repeating the send. Never claim that a vector store or Redis lock has removed that uncertainty.
Retrieval can begin inside pgvector. If it moves to Qdrant, Pinecone or Weaviate, mirror only the derived chunks and their IDs. The application binds the client's authorised namespace, tenant or payload filter; retrieved material is context, not a set of trusted commands. Redis, if added, caches a run's temporary results with an expiry. Removing that cache should slow the agent, not erase its accepted decisions.
For the two Northstar accounts, a semantic query may find both names interesting. The authenticated client boundary must exclude the other client's documents before they reach the model. A prompt asking the model to ignore the wrong client is too late.
Price the actual workload
Our fictional workload has 100,000 document chunks, 1,536-dimensional float32 embeddings, 20,000 retrieval queries monthly and ten clients. Assume 4 GB of total PostgreSQL disk for the pricing illustration; that includes more than raw vector bytes and is a sizing assumption, not measured index capacity.
The raw vector arithmetic is:
100,000 × 1,536 × 4 bytes = 614,400,000 bytes, about 0.614 GB in decimal units. Text averaging 2,000 bytes per chunk adds another 0.2 GB. Indexes, metadata, replicas and database overhead are additional. At one million chunks, raw vectors alone become 6.144 GB. Do not use raw vector size as a complete memory or disk quote.
| Configuration | Supported pricing calculation | What the figure means |
|---|---|---|
| Supabase Pro, one Micro | $25 plan + $10 compute − $10 credit | $25/month base, within included allowances and only if that size serves the workload |
| Supabase Pro, one Medium | $25 + $60 − $10 | $75/month base, before extras |
| Neon Launch, intermittent illustration | 88 CU-hours × $0.106 + 4 GB × $0.35 + 2 GB history × $0.20 | $11.13/month for these specified components |
| Neon Launch, always-on illustration | 365 CU-hours × $0.106 + same disk/history | $40.49/month for these specified components |
| PostgreSQL plus Pinecone Standard | Relational base + max($50, metered Pinecone usage) | At least $75/month using the $25 relational base |
| PostgreSQL plus Weaviate Flex | Relational base + Flex configuration from $45 | At least $70/month using the $25 relational base |
| PostgreSQL plus Qdrant Standard | Relational base + selected resource-priced cluster | Cluster basis; no fabricated fixed total |
The intermittent Neon example assumes 0.5 CU active for eight hours on 22 days: 88 CU-hours. The always-on case assumes 0.5 CU for a 730-hour month: 365 CU-hours. History uses the assumed 2 GB of chargeable changed data, not an extra copy automatically inferred from the 4 GB database. Network, additional branches, snapshots and other services are outside those component totals. The provider's “typical spend” figure is not a fixed plan price.
The separate-vector totals retain PostgreSQL because the workload still needs durable state. They are floors, not a claim that 100,000 vectors and 20,000 queries always fit the minimum configuration. Embedding generation, agent model calls and reranking are separate in every architecture.
Add operator time. If the single-database design needs one hour monthly at $75, the Micro illustration becomes $100 and Medium $150. If an external index requires two hours for synchronisation and operations, the minimum PostgreSQL-plus-Pinecone allocation becomes $225 and the Weaviate allocation $220, before extra traffic or model costs.
Those are workload assumptions, not measured labour differences between vendors. They make the purchase question concrete: will a separate retrieval service save enough time, improve useful retrieval or meet a scaling need to outweigh the extra service and integration? Starting with one database is sensible until that answer becomes yes.
Freshness, deletion and recovery
Suppose a website says a contact is the procurement lead, but the CRM now marks them as having left. The agent can cite the dated webpage as historical evidence; it must not overwrite the current contact record from the old summary. Keep source date and current business status available to the action step.
For deletion, track a source record through its document versions, chunks, vector IDs, summaries and cache keys. Mark it unavailable in the authoritative store immediately, then remove the derived copies. Until removal propagates, a final source-status check prevents a stale retrieved chunk from being used. Apply the same approach to suppression: the sending decision reads current durable state.
Backups have a retention policy, not an instant per-row erase guarantee. A restored backup should replay the subsequent deletion/suppression record before serving traffic; otherwise recovery can revive material deliberately removed from the live system. Retain only the minimal removal marker needed for that process.
For retrieval quality, build a labelled question set from the actual research task. Include exact company-name questions, product terms, outdated-source questions and the two-client Northstar case. Compare useful answers and source support against an exact-search baseline before paying for larger approximate indexes. These are proposed evaluation cases, not claimed benchmark results.
Which system should you choose?
Start with PostgreSQL and pgvector when a technical GTM team needs account state, run history and a moderate evidence corpus. Managed PostgreSQL can be inexpensive enough that operating a server solely to avoid the subscription is a poor saving.
Add Redis for a clear temporary-state or coordination problem. Add Qdrant when specialised filtered retrieval and deployment control matter. Choose Pinecone when managed vector operations and the relevant paid-plan features reduce the team's burden. Choose Weaviate when hybrid retrieval and its tenant structure match the research application.
Keep approval, suppression and external-action identity in durable records regardless of the retrieval choice. A persistent agent should remember the right information with its authority and date, not simply retrieve more text. Our signal-based outbound playbook connects the resulting account evidence to a practical sales decision.
FAQ
Can a vector database replace the CRM?
It can retrieve documents and similar records, but CRM fields need deterministic identity, business rules and accepted updates. Keep the CRM authoritative for its stages and contacts, and use retrieval to provide context. A high similarity score is not permission to change a business field.
Is Redis unsuitable for durable data?
It can persist data when configured for the requirement. The question is the selected persistence, replication, eviction and recovery behaviour. For this GTM architecture, keeping approvals and suppression in the relational ledger is simpler than asking an expiring working-state layer to serve every role.
How many vectors make pgvector too small?
There is no useful universal cutoff. Dimensions, filtering, index type, concurrency, update frequency and hardware all affect the result. Begin with the actual question set and latency requirement. The byte calculation sizes raw embeddings; it does not certify throughput or index memory.
Do namespaces or tenant fields provide authentication?
They organise and scope stored data. The application must determine which tenant the authenticated caller may use. Do not accept a tenant identifier merely because an agent supplied it. PostgreSQL roles, authorised service requests and provider controls should support the same boundary.
Should every agent conversation be saved forever?
No. Keep durable decisions and useful evidence with a defined purpose and date. Temporary tool output and conversational state can expire. Summaries can be regenerated. Saving everything increases stale retrieval and deletion work without necessarily improving the next sales decision.
Is the cheapest vector plan the cheapest agent architecture?
Only if it supplies the required features and capacity. You still need the durable record system, model calls, source documents and integration work. The worked comparison retains those dependencies rather than replacing a transactional database with a $20 or $45 headline.





