OpenAI API integration

A model reads incoming emails and PDFs, then writes the result into the tools you already use.

  • 15Meetings booked by the voice agent in week oneE-Ihracat Turkiye reports
  • $8Kin sales from week-one voice‑agent meetingsE-Ihracat Turkiye reports
Supplier delay email at Quillan IndustrialExample
  1. Supplier email about a late shipmentTrigger · Starts the run
  2. Strip signature and attachmentsn8n · Only the text the model needs is sent
  3. Read the email into schema fieldsOpenAI · PO-7731, new ship date Oct 14
  4. Validate against the open PONetSuite · PO found, both lines match
  5. Update the expected receipt dateNetSuite · Oct 14 set on both lines
  6. Flag the customer orders affectedTeams · A rep approves the delay notice
PO dates updated before anyone read the email.

Why most model API integrations stall after the demo

  • Answers the business cannot stand behind

    Asked about a return policy or an order status it was never given, the model produces a plausible answer anyway.

  • No test set, so every prompt change is a gamble

    Someone tweaks the prompt to fix one bad case and quietly breaks three others.

  • A bill nobody can explain

    Usage is billed per token.

  • Keys and data in the wrong hands

    The key sits in a contractor's personal account or is pasted into a spreadsheet.

  • Start with the one that costs the most.

    On a free 30-minute call we go through your week and agree which of these to fix first.

How an OpenAI API integration project runs

  1. Pick the step

    We map the process and agree the one step where a model saves the most time or money, with a measure such as minutes per ticket or documents processed per day.

  2. Collect real cases

    Your team supplies a sample of real inputs and correct outputs.

  3. Test models and prompts

    Candidate models and prompt designs run against the set.

  4. Build the integration

    Trigger, cleanup, model call with schema, validation, human review queue and write-back, in your accounts with your keys, with logging on every call.

  5. Run in shadow, then live

    The integration runs beside your team first so outputs can be compared.

  6. Hand over and monitor

    You get documentation, the evaluation set and a cost per task baseline.

★★★★★

Benian Technologies was a great investment. I wanted him to connect my crm to a automatic calling agent. He built so many more connections than I expected. Takes notes of the calls, and the agent speaks the way we would speak to customers. After our discovery and strategy call we established the roadmap and he delivered with flying colors!🚀💪👍

Derin GocekOwner, Deep Sea MediaGoogle review · April 2026

Questions we get asked

How do I integrate the OpenAI API into my business software?

Create an API account in your company's name, then connect it through your automation tool or a small service that calls the API when an event happens, such as a new email or form. Ask for structured JSON output, validate it, and write the result into the system your team already uses. The work beyond that first call is handling bad inputs, reviews and cost.

Should we use OpenAI or Claude?

Test both on a sample of your own inputs, because general rankings rarely predict performance on your documents and rules. Both offer per token pricing, several model sizes, structured output and tool calling. Build so the provider can be swapped, and the choice stops being a long‑term bet.

Can I build an AI agent in Microsoft Copilot Studio instead?

Yes, and if your team lives in Teams and the agent mainly answers questions from SharePoint and other Microsoft 365 content, it is often the faster route. Custom builds make sense when the agent must act in systems outside Microsoft, needs logic Copilot Studio cannot express or must pass your own test set before each change ships.

How do we stop the model from making things up?

Give it the relevant records in each request, require it to cite which passage it used and tell it to say it does not know when nothing relevant was provided. Validate structured outputs in code. Keep a human approval step on anything that commits money or makes a promise to a customer, and measure the error rate on a fixed test set.

More questions
Who owns the API keys and data?

In a Benian build, your business does. The provider account, billing and keys are in your name, stored in your secrets store or automation tool, and our access is revocable by you. Prompts, evaluation sets and logs are delivered to you as part of the work.

What drives the running cost of an LLM integration?

Tokens per task, task volume and the model chosen. Long prompts, whole documents sent when a paragraph would do and a large model used for simple sorting are the usual causes of a high bill. Logging cost per task from day one shows which step to trim.

Do we need LangGraph to build an agent?

Usually not. Most business integrations are a fixed sequence with one or two model calls, which plain code or an n8n workflow handles well. LangGraph helps when an agent needs loops, saved state and pauses for human approval across many steps.

Read the full guide9 min read

An OpenAI API integration is worth building when a model call replaces a step your team does by hand every day, such as reading an inbound email, pulling the fields out of a PDF or drafting a reply, and the output lands in the system that already runs the work. The first API call takes an afternoon. Making it right on the five hundredth messy email, at a cost you can predict, is the actual project.

Benian Technologies is an AI implementation partner. We find the step in your process that costs time or revenue, then build the model call into it with the OpenAI API, the Claude API or Azure OpenAI, whichever fits the job. The build runs on API keys and accounts your business holds, so the usage bill, the logs and the prompts are yours.

This page covers where a model belongs in a process, how to choose between providers, how to keep the output structured and grounded, how to test it, what drives the running cost, and when Copilot Studio or a framework like LangGraph makes more sense than custom code.

Why most model API integrations stall after the demo

Free text where software needs fields

The model returns a friendly paragraph, and the next step in the workflow needs a customer ID, an amount and a category. Someone ends up copying values out of the reply, which is the job the integration was meant to remove.

Answers the business cannot stand behind

Asked about a return policy or an order status it was never given, the model produces a plausible answer anyway. Without grounding in your own records and a rule to say it does not know, every reply needs a human check.

No test set, so every prompt change is a gamble

Someone tweaks the prompt to fix one bad case and quietly breaks three others. Nobody notices until a customer does, because there was never a list of real cases to rerun.

A bill nobody can explain

Usage is billed per token. Long prompts, full documents pasted into every call and the largest model used for simple sorting all multiply cost, and the invoice arrives before anyone has measured cost per task.

Keys and data in the wrong hands

The key sits in a contractor's personal account or is pasted into a spreadsheet. When the contractor leaves, the integration stops, or worse, keeps running on an account you cannot see.

Where an OpenAI API integration belongs in a business process

A model is good at one kind of step: turning unstructured input into a decision or a draft. Reading a supplier email and deciding which order it refers to. Classifying a support ticket by urgency and product. Extracting line items from an invoice PDF. Drafting a first reply from your own policy text. Those steps sit in the middle of a workflow, so the integration is mostly plumbing around one call.

Steps that follow fixed rules do not need a model. If a field always maps to the same place, a plain automation in n8n or code is cheaper, faster and never makes things up. We often remove the model from a proposed design because a rule does the job. That is the right answer when it is true.

A typical openai automation we build looks like this: a trigger such as a new email or form, a cleanup step that strips signatures and attachments the model does not need, one model call that returns structured fields, a validation step, then a write into your CRM, help desk or accounting tool. Anything below a confidence threshold, or touching money or a customer commitment, goes to a person for approval first.

  • Good fits: intake triage, document field extraction, reply drafts for review, call and meeting summaries into the CRM, matching messy records.
  • Poor fits: calculations, anything with a single correct lookup, legal or medical judgment without a qualified human signing off.

Choosing a model: OpenAI, Claude or Azure OpenAI

OpenAI and Anthropic, the maker of Claude, both sell API access billed per token, with several model sizes at different speeds and costs. Both support structured output and tool calling. In practice the difference shows up on your own cases: one handles your long contracts better, the other follows your formatting rules more reliably. We test the candidates on a sample of your real inputs before committing, and we design the integration so switching later is a configuration change, not a rebuild.

Azure OpenAI runs OpenAI models inside a Microsoft Azure subscription. It suits companies that already buy through Azure, want usage on their existing Microsoft agreement or need the deployment under their own Azure governance and network controls. The trade is more setup: resource provisioning, model deployments and quota requests in your tenant. If you are building an azure ai chatbot or running ai agents on Azure that your own IT team will manage, that setup is worth it. If not, the direct API is simpler.

Model choice is rarely the decision that sinks a project. Unclear inputs, no test set and no human review step sink far more.

Structured outputs and tool use

Both APIs can be told to return JSON that matches a schema you define: the exact fields, their types and which values are allowed. That turns the model from a writer into a component. The workflow reads order_id and refund_reason directly, and a response that fails the schema is caught and retried or routed to a person instead of flowing downstream.

Tool use, also called function calling, lets the model ask your code to do something: look up an order, check calendar availability, fetch a customer record. The model proposes the call; your code decides whether to run it. We keep tools narrow and read only by default. A tool that writes, refunds or sends gets an approval step or a hard limit in code, because a prompt is not a permission system and text inside an email can try to steer the model.

Grounding answers in your own data

The model only knows what is in the request. To answer from your price sheet, your policies or a customer's history, the integration has to fetch the right records and pass them in. For a small, stable set of documents that can be as simple as including them. For larger libraries it means retrieval: index the documents, find the passages relevant to each question and send only those.

Retrieval fails in predictable ways. Outdated documents get indexed beside current ones. Scanned PDFs come through as noise. The right passage exists but is phrased differently from the question. We clean the source set before indexing, require the model to cite which passage it used and instruct it to say it does not know when nothing relevant comes back. A reviewer can then check the citation in seconds instead of rereading the source.

Evaluations before launch and after every change

Before an integration goes live we build an evaluation set: real inputs from your business with the output a competent employee would produce. A few dozen to a few hundred cases is a common starting size, weighted toward the awkward ones, such as the email in two languages, the invoice with a missing total, the customer who asks two questions at once.

Every prompt edit, model change or new tool reruns that set and reports what got better and what got worse. Field extraction is scored by exact match. Drafts are scored against a written rubric, partly by a person. After launch, cases your team corrects are added to the set, so the test gets harder as the system meets reality.

Cost, latency and rate limits

Running cost is driven by tokens: how much text goes in and comes out per task, times the number of tasks, times the rate for the model chosen. The levers are concrete. Send only the text the step needs. Use a smaller model for sorting and extraction and save the larger one for reasoning-heavy steps. Cache repeated instructions where the provider supports it. Batch work that can wait overnight.

Latency matters when a person is waiting, on a chat widget or a phone call. Background work such as document processing can take longer and run in batches. Each provider also enforces rate limits per account, so a backlog of a few thousand documents needs queueing and retry with backoff, or the job fails halfway. We log tokens and duration for every call so you can see cost per task, not only the monthly total.

Agent frameworks: LangGraph, Copilot Studio or plain code

Most business integrations are a fixed sequence with one or two model calls, and plain code or an n8n workflow handles that well. An agent, where the model decides which tools to call and in what order, is justified when the path through the task genuinely varies case by case.

LangGraph is an open source framework for building ai agents in LangGraph style graphs: each step is a node, and the framework keeps state, supports loops and lets you pause for human approval. It earns its place in multi-step agents that need checkpoints and recovery. It adds a dependency and a learning curve your next developer will also need.

Copilot Studio is Microsoft's tool to build ai agent in Copilot setups that live inside Teams and Microsoft 365, connected to SharePoint and other Microsoft data. If your staff work in Teams and the agent mainly answers questions from internal documents, it is often the faster route, and you may not need us. Custom microsoft ai agent development makes sense when the agent must act in systems outside Microsoft, needs custom logic or must be tested against your own evaluation set.

Data privacy, keys and who owns the account

The API account, the billing and the keys sit in your organization's name. We work through access you grant and can revoke. Keys live in a secrets store or the credential vault of your automation tool, never in prompts, spreadsheets or chat messages, and each integration gets its own key so one can be rotated without breaking the rest.

Read the provider's current data use and retention terms for the account type you hold, and decide which data may leave your systems before anything is sent. We strip fields the model does not need, such as full card numbers or government IDs, before the call. If a data category cannot go to a third party under your contracts, that step stays out of the model.

When not to hire anyone for this

If the job is one person drafting emails faster, a ChatGPT or Claude subscription with good instructions is enough. If what your team lacks is the skill to use these tools well, training comes before integration. E-Ihracat Turkiye, an e-commerce education and services company, bought Claude enablement from Benian: 20 hours of hands-on training on their real accounts, which the client credits with its revenue change shown on this page. That was training, not an API build, and for many teams it is the right first step.

How an OpenAI API integration project runs

  1. Pick the step. We map the process and agree the one step where a model saves the most time or money, with a measure such as minutes per ticket or documents processed per day.
  2. Collect real cases. Your team supplies a sample of real inputs and correct outputs. This becomes the evaluation set and usually exposes edge cases nobody had written down.
  3. Test models and prompts. Candidate models and prompt designs run against the set. We pick on accuracy, cost per task and speed, and share the scores.
  4. Build the integration. Trigger, cleanup, model call with schema, validation, human review queue and write-back, in your accounts with your keys, with logging on every call.
  5. Run in shadow, then live. The integration runs beside your team first so outputs can be compared. It goes live once results match the agreed threshold, with approval steps kept on risky actions.
  6. Hand over and monitor. You get documentation, the evaluation set and a cost per task baseline. Corrections from your team feed the next round of tests.

Stop copying fields out of emails and PDFs.

A free 30-minute call about your business, your systems and what you want to build.